kaxil opened a new pull request, #74266: URL: https://github.com/apache/airflow/pull/74266
pydantic-ai 2.53.0 added `SystemOneModel` (pydantic/pydantic-ai#8942), which runs any decision model served over the same `POST /v1/systemone` API as TypeSafe's Jev: decision models in Ollama 0.35+, AWS's [Strands Decider](https://github.com/strands-labs/strands-decider), [Kev](https://github.com/jaredpalmer/kev), CLM, Laya. The Common AI docs still said Jev was the only decision model pydantic-ai supports. This needs no code change. `PydanticAIHook` resolves an unknown prefix through its generic path, `infer_provider_class(prefix)(api_key=conn.password, base_url=conn.host)`, and `SystemOneProvider` takes exactly those two arguments. A `pydanticai` connection with the server URL in **Host** and `system-one:<model>` as its **Model** is the whole integration. The end-to-end run below confirms it. So this PR is docs and examples: - `classifier_models.rst` becomes `decision_models.rst` (redirect added). "Decision model" is the term pydantic-ai, Strands and Cloudflare all use now, and the operators already call the gate `decision_policy`. Setup covers both backends, `typesafe:` and `system-one:`. - `example_classifier_model.py` becomes `example_decision_model.py`, and its Dags become `example_decision_model_branch` / `example_decision_model_confidence`. These Dags, `example_llm_branch_decision_policy` and the decision-model retry policies now read a `decision_default` connection and take the model from it rather than hard-coding `model_id="typesafe:jev-1.13.0"`, so the same Dag runs on either backend. - Docstrings and the retry-policy log hint stop saying "(TypeSafe's)". ## Design rationale **The examples now describe every option.** The first end-to-end run failed with a 422: for an option with no description, pydantic-ai sends `criteria: {option: null}`, and Strands Decider 0.1.0 accepts only string criteria. It also requires question text, which an operator with the default empty `system_prompt` does not send. The branch example now passes `branches=`, and the classify example uses an `Enum` with `UseEnumMemberDocstrings` instead of a bare `Literal`. The decision models page documents both requirements, since they are the first thing a Strands Decider user would hit, and the `branches` docstring on `LLMBranchOperator` says the same. **A `system-one:` model's option cap is the server's.** pydantic-ai refuses a question over Jev's cap before sending it. A model reached through a connection carries no `DecisionModelProfile`, so a question over a System One server's cap (Ollama's is 26) comes back as an HTTP error instead. The page says so. **Cloudflare's Clef is not named.** Its changelog says it "follows the System One API", but Workers AI serves it at `/ai/run`. I have not confirmed that `SystemOneModel` can reach it, so the docs don't claim it. ## Gotchas - `example_decision_model.py` imports `UseEnumMemberDocstrings`, which shipped in pydantic-ai-slim 2.46. The provider floor stays at 2.33. The example-Dag import test already skips under lowest dependencies, and the page and the example state the versions needed (2.46 for the examples, 2.53 for `system-one:`). On an older pydantic-ai, a `system-one:` connection fails with the hook's existing "is not a provider pydantic-ai recognizes" error. - Example Dag ids and the examples' connection id change (`jev_default` to `decision_default`). The retry-policy example keeps its note that its confidence bars were calibrated on `jev-1.13.0` and need measuring again for another model. - These examples were not re-run against Jev for this PR. For Jev, the change to them is the connection id and the added option descriptions. ## End-to-end run Airflow `main` under breeze, sqlite, pydantic-ai-slim 2.53.0. Strands Decider 2B (`StrandsAgents/strands-decider-2B-hobson-v19`) served locally by `strands-decider serve`. The three Dags ran unmodified from the provider's example-Dag bundle. The classify step was re-run through `PydanticAIHook` after the example's `Enum` got its class docstring: still `resource`, confidence 0.61. | Dag | Result | |---|---| | `example_llm_branch_decision_policy` | Picked `rerun` at confidence 0.74, above the 0.6 bar. `page_oncall` and `ignore` skipped | | `example_decision_model_branch` | Picked `grant_bucket_write`. `restore_deleted_bucket` and `wait_and_retry` skipped | | `example_decision_model_confidence` | Classified `resource` at 0.63, and `act` filed it for review (0.5 to 0.8 band) | The connection, with the server URL in Host and the `system-one:` model:   The `decision` XCom from `example_llm_branch_decision_policy`, with the model, confidence and probabilities Strands Decider returned:  `example_decision_model_confidence`. The red runs in the grid are the earlier attempts: the previous example's undescribed options (the 422 above), plus a crash of the local model server.  ## Follow-ups `ToolCallJudge` (pydantic-ai-harness 0.53) is a natural fit for unattended agents, but it builds its judge model when the Dag file is parsed, so today its credentials can only come from environment variables, not an Airflow connection. That gets its own PR. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
