This is an automated email from the ASF dual-hosted git repository.

kaxil pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/airflow.git


The following commit(s) were added to refs/heads/main by this push:
     new 758ebbd35f4 Document decision models served over the System One API in 
common.ai (#74266)
758ebbd35f4 is described below

commit 758ebbd35f4ad963d29058a326c719626d9f6000
Author: Kaxil Naik <[email protected]>
AuthorDate: Mon Oct 5 21:00:57 2026 +0100

    Document decision models served over the System One API in common.ai 
(#74266)
    
    pydantic-ai 2.53.0 added SystemOneModel, which runs any decision model
    behind the POST /v1/systemone API Jev uses (Ollama, Strands Decider, Kev
    and others). PydanticAIHook already builds it from a pydanticai connection
    through its generic provider path, so this is docs and examples only.
    
    Rename the "Classifier models" page to "Decision models" with a redirect,
    cover both the typesafe: and system-one: backends in setup, and make the
    decision-model examples read a decision_default connection with every
    option described, so they run unchanged on either backend.
    
    Co-authored-by: Claude <[email protected]>
---
 docs/spelling_wordlist.txt                         |   1 +
 providers/common/ai/docs/approval_gates.rst        |   4 +-
 .../{classifier_models.rst => decision_models.rst} | 158 +++++++++++++++------
 providers/common/ai/docs/examples.rst              |   6 +-
 providers/common/ai/docs/installation.rst          |   5 +-
 providers/common/ai/docs/model_providers.rst       |  12 +-
 providers/common/ai/docs/operators/llm_branch.rst  |  14 +-
 providers/common/ai/docs/redirects.txt             |   1 +
 providers/common/ai/docs/retry_policies.rst        |  15 +-
 providers/common/ai/docs/self_hosted_models.rst    |   3 +
 providers/common/ai/docs/stability.rst             |   2 +-
 providers/common/ai/docs/use_cases/index.rst       |   3 +-
 .../ai/docs/use_cases/route_pipeline_failures.rst  |  25 ++--
 ...assifier_model.py => example_decision_model.py} |  79 +++++++----
 .../common/ai/example_dags/example_llm_branch.py   |   5 +-
 .../ai/example_dags/example_llm_retry_policy.py    |  22 +--
 .../airflow/providers/common/ai/operators/llm.py   |   8 +-
 .../providers/common/ai/operators/llm_branch.py    |  14 +-
 .../providers/common/ai/policies/decision.py       |   4 +-
 .../airflow/providers/common/ai/policies/retry.py  |  19 ++-
 .../airflow/providers/common/ai/utils/decision.py  |  12 +-
 .../unit/common/ai/operators/test_llm_branch.py    |   2 +-
 .../ai/tests/unit/common/ai/policies/test_retry.py |   2 +-
 23 files changed, 260 insertions(+), 156 deletions(-)

diff --git a/docs/spelling_wordlist.txt b/docs/spelling_wordlist.txt
index 6db426432a6..325b68d6c90 100644
--- a/docs/spelling_wordlist.txt
+++ b/docs/spelling_wordlist.txt
@@ -972,6 +972,7 @@ kerberized
 Kerberos
 kerberos
 KerberosClient
+Kev
 keycloak
 Keyfile
 keyfile
diff --git a/providers/common/ai/docs/approval_gates.rst 
b/providers/common/ai/docs/approval_gates.rst
index 39b16ea09b5..c40eee8e2e9 100644
--- a/providers/common/ai/docs/approval_gates.rst
+++ b/providers/common/ai/docs/approval_gates.rst
@@ -90,8 +90,8 @@ Reviewing uncertain output
     Experimental: this can change or be removed in a minor release of this 
provider.
     See :ref:`howto/stability`.
 
-A classifier model such as TypeSafe's reports a confidence for every field
-of a structured output, in ``provider_details`` on the model response. It is a
+A decision model reports a confidence for every field of a structured output, 
in
+``provider_details`` on the model response. It is a
 summary of how concentrated the model's probability distribution was, not the
 probability that the field is right. 
``decision_policy=DecisionPolicy(min_confidence=0.7)``
 (import ``DecisionPolicy`` from ``airflow.providers.common.ai.operators.llm``)
diff --git a/providers/common/ai/docs/classifier_models.rst 
b/providers/common/ai/docs/decision_models.rst
similarity index 52%
rename from providers/common/ai/docs/classifier_models.rst
rename to providers/common/ai/docs/decision_models.rst
index 2f003806e56..3f32da16aa2 100644
--- a/providers/common/ai/docs/classifier_models.rst
+++ b/providers/common/ai/docs/decision_models.rst
@@ -15,53 +15,110 @@
     specific language governing permissions and limitations
     under the License.
 
-Classifier models
-=================
+Decision models
+===============
 
 .. note::
 
     Experimental: this can change or be removed in a minor release of this 
provider.
     See :ref:`howto/stability`.
 
-Every other model in this provider writes text. A classifier model does not: 
you give it
+Every other model in this provider writes text. A decision model does not: you 
give it
 some text and a typed question, and it answers with a value from a set you 
named in
 advance, plus a confidence. Ask it for a string and the request is refused 
before it
 leaves your process.
 
-`TypeSafe <https://typesafe.ai>`__'s Jev is the one pydantic-ai supports, as 
the
-``typesafe:`` provider. Nothing in this provider is specific to it -- it 
arrives through the same
+pydantic-ai reaches decision models through two model prefixes. Nothing in 
this provider is
+specific to either: both arrive through the same
 :class:`~airflow.providers.common.ai.hooks.pydantic_ai.PydanticAIHook` as 
every other
 model, so a model id is the whole integration.
 
+.. list-table::
+   :header-rows: 1
+   :widths: 18 32 50
+
+   * - Prefix
+     - Needs
+     - Runs
+   * - ``typesafe:``
+     - ``pydantic-ai-slim`` 2.45.0 or later (2.46.0 for the examples on this 
page, which
+       describe options with ``UseEnumMemberDocstrings``) and the provider's 
``typesafe`` extra
+     - `TypeSafe <https://typesafe.ai>`__'s hosted Jev.
+   * - ``system-one:``
+     - ``pydantic-ai-slim`` 2.53.0 or later, no extra
+     - Any server that answers the same ``POST /v1/systemone`` API as Jev, 
hosted or your
+       own. `Ollama <https://docs.ollama.com/capabilities/decision>`__ 0.35 
and later serves
+       it for the decision models it runs, ``strands-decider serve`` serves 
AWS's
+       `Strands Decider <https://github.com/strands-labs/strands-decider>`__, 
and
+       `Kev <https://github.com/jaredpalmer/kev>`__ is an open-weight model 
built to serve
+       it. `pydantic-ai's System One page 
<https://pydantic.dev/docs/ai/models/system-one/>`__
+       lists more.
+
+Both prefixes need a newer ``pydantic-ai-slim`` than the floor the provider's 
other extras
+set, so check the release you have installed.
+
 Setup
 -----
 
-1. Install the extra:
+The examples on this page use a connection named ``decision_default``. Create 
it for the
+backend you run.
+
+**TypeSafe Jev**
+
+1. Install the extra, which adds the TypeSafe SDK:
 
    .. code-block:: bash
 
        pip install 'apache-airflow-providers-common-ai[typesafe]'
 
-   The extra installs the TypeSafe SDK. The ``typesafe:`` model adapter is 
part of
-   pydantic-ai itself from ``pydantic-ai-slim`` 2.45.0, which is newer than 
the floor the
-   provider's other extras set, so check that release or later is installed.
-
 2. Create a connection (``Admin > Connections``):
 
-   - **Connection Id**: ``jev_default``
+   - **Connection Id**: ``decision_default``
    - **Connection Type**: ``Pydantic AI``
    - **Password**: your TypeSafe API key, from your `TypeSafe account 
<https://typesafe.ai>`__
    - **Extra**: ``{"model": "typesafe:jev-1.13.0"}``
 
-Leave **Host** empty unless you are pointing at a proxy; the provider defaults 
to
-TypeSafe's own endpoint.
+   Leave **Host** empty unless you are pointing at a proxy; the provider 
defaults to
+   TypeSafe's own endpoint.
+
+**A System One server**
 
-Pin the version rather than using ``jev-latest``. A threshold you tuned 
against one
-release is not guaranteed to mean the same thing after the next one, and 
``jev-latest``
-moves under you.
+1. Start the server, or note the URL of a hosted one. For example, to serve 
Strands Decider
+   on your own machine:
 
-When a classifier model is the right choice
--------------------------------------------
+   .. code-block:: bash
+
+       pip install strands-decider
+       strands-decider serve StrandsAgents/strands-decider-2B-hobson-v19 
--port 8000
+
+2. Create a connection (``Admin > Connections``):
+
+   - **Connection Id**: ``decision_default``
+   - **Connection Type**: ``Pydantic AI``
+   - **Host**: the server's URL, such as ``http://decider.internal:8000``, 
with or without a
+     trailing ``/v1``
+   - **Password**: the server's API key, sent as a bearer token. Leave it 
empty for a server
+     that takes none, such as a local Ollama.
+   - **Extra**: ``{"model": "system-one:strands-decider-2B-hobson-v19"}``
+
+   The name after ``system-one:`` is sent to the server as the model to answer 
with. A
+   server that runs several, such as Ollama, picks one by it.
+
+3. Describe every option you ask about, and give the question its text. 
Servers differ in
+   what they accept, and Strands Decider 0.1.0 refuses, with an HTTP 422 that 
fails the task,
+   a question that has an option without a description or that has no question 
text at all.
+   Describe each branch in ``branches`` and each member of an ``Enum`` through 
a docstring;
+   the members of a bare ``Literal`` have no descriptions, so use a described 
``Enum``
+   instead, as the worked example below does. The question text is the 
operator's
+   ``system_prompt`` or the agent's ``instructions``, which default to empty.
+
+Pin the model and its version, as in ``typesafe:jev-1.13.0`` rather than
+``typesafe:jev-latest``. A threshold you tuned against one model, or one 
release of it, is
+not guaranteed to mean the same thing on another, so measure it again when you 
change
+either.
+
+When a decision model is the right choice
+-----------------------------------------
 
 All four of these have to hold.
 
@@ -70,11 +127,14 @@ All four of these have to hold.
 ``str`` field is refused, so anything that writes a summary, a query, a 
migration, or a
 reply to a person is out.
 
-**You can name the options up front, and there are at most 255.** Downstream 
task ids,
-error categories, severity levels, environments, a repository list. If the set 
is open, or
-is discovered at run time and could grow past the cap, this is the wrong tool. 
The cap is
-per question, so tools attached to an agent get their own 255, with the output 
type as one
-option in that question.
+**You can name the options up front, and there are few enough for the model.** 
Downstream
+task ids, error categories, severity levels, environments, a repository list. 
If the set is
+open, or is discovered at run time and could grow past the cap, this is the 
wrong tool. Each
+model has its own cap per question: Jev takes 255 options, and Ollama takes 
26. Tools
+attached to an agent count as options in a question of their own, with the 
output type as
+one more. pydantic-ai refuses a question over Jev's cap before sending it; a 
``system-one:``
+model's cap is the server's, so a question over it comes back as an HTTP error 
from the
+server instead.
 
 **The decision is on a path where latency or cost is the constraint.** One 
decision per
 Dag run rarely justifies changing models. One per row, per file, or per 
retrieved document
@@ -88,7 +148,7 @@ itself. If you would not do anything different with a 
confidence of 0.55 than wi
 that is a sign a general-purpose model is fine here.
 
 And one case where the answer is neither: if a deterministic rule already 
sorts the input
-correctly, use the rule. A classifier model is cheap, not free, and a rule you 
can read is
+correctly, use the rule. A decision model is cheap, not free, and a rule you 
can read is
 worth more than a probability you have to calibrate.
 
 Where it fits in this provider
@@ -104,9 +164,10 @@ Where it fits in this provider
    * - 
:class:`~airflow.providers.common.ai.operators.llm_branch.LLMBranchOperator`
      - Yes, with a caveat
      - The downstream task ids are already presented to the model as a 
constrained set of
-       choices, which is exactly the shape a classifier model answers. Setting
-       ``model_id`` is the only change, as long as there are two or more 
downstream tasks
-       -- a one-option pick is refused. Describe each branch in ``branches`` 
and set a
+       choices, which is exactly the shape a decision model answers. Pointing
+       ``llm_conn_id`` (or ``model_id``) at one is the only change, as long as 
there are two
+       or more downstream tasks -- a one-option pick is refused. Describe each 
branch in
+       ``branches`` and set a
        ``decision_policy`` so an unsure pick goes to a person instead of 
branching; see
        :doc:`operators/llm_branch`.
    * - :class:`~airflow.providers.common.ai.operators.llm.LLMOperator` /
@@ -115,7 +176,7 @@ Where it fits in this provider
      - Yes
      - A ``Literal``, ``Enum``, ``bool`` or bounded number works. Describe the 
field, which
        becomes the question, and describe each option, which is what tells 
them apart. An
-       option with no description is read from its name alone.
+       option with no description is read from its name alone, and some 
servers refuse it.
    * - :doc:`ClassifierRetryPolicy <retry_policies>`
      - Yes
      - The model names one of the policy's ``categories`` and nothing else; 
retry or
@@ -123,12 +184,12 @@ Where it fits in this provider
        worker. Set ``min_confidence`` and an unsure answer goes to 
``fallback_policy``
        (typically an ``LLMRetryPolicy`` on a text model), then 
``fallback_rules``, then
        the task's own retry behaviour, instead of ending the task on the 
model's say-so.
-       ``LLMRetryPolicy`` itself asks for free text, which a classifier model 
refuses.
+       ``LLMRetryPolicy`` itself asks for free text, which a decision model 
refuses.
        This is the surface where the model's speed and price matter most: it 
runs on
        every task failure.
    * - Agents with toolsets
      - Partly
-     - Which tool the text calls for is itself a pick, so a classifier model 
can make it.
+     - Which tool the text calls for is itself a pick, so a decision model can 
make it.
        What it cannot write is a tool's arguments. A tool taking none it calls 
itself; one
        taking arguments raises ``ToolCallProposed`` after the request, which 
is a
        ``ModelAPIError`` rather than a refusal, so ``FallbackModel`` hands 
those requests to
@@ -155,21 +216,35 @@ transcript when that is enabled, and a hook-level call 
has it on the result:
 
 .. code-block:: python
 
+    from enum import Enum
+
+    from pydantic_ai import UseEnumMemberDocstrings
+
     from airflow.providers.common.ai.hooks.pydantic_ai import PydanticAIHook
     from airflow.sdk import task
-    from typing import Literal
+
+
+    class FailureCause(UseEnumMemberDocstrings, str, Enum):
+        transient = "transient"
+        """A fault that clears by itself: a timeout, throttling, a dropped 
connection."""
+
+        resource = "resource"
+        """A dependency is down or unreachable and needs fixing before a retry 
can work."""
+
+        permanent = "permanent"
+        """A bug or bad input that fails the same way however often it is 
retried."""
 
 
     @task
     def triage(log_line: str) -> dict:
-        agent = PydanticAIHook(llm_conn_id="jev_default").create_agent(
-            output_type=Literal["transient", "resource", "permanent"],
+        agent = PydanticAIHook(llm_conn_id="decision_default").create_agent(
+            output_type=FailureCause,
             instructions="Classify why this Airflow task failed.",
         )
         result = agent.run_sync(log_line)
         details = result.response.provider_details or {}
         confidence = (details.get("confidence") or {}).get("response")
-        return {"category": result.output, "confidence": confidence}
+        return {"category": result.output.value, "confidence": confidence}
 
 Then branch on the returned confidence in a downstream task, so an unsure 
answer escalates
 instead of acting. Use a higher bar for acting automatically than for flagging 
something
@@ -179,21 +254,22 @@ shape for the decision.
 Worked example
 --------------
 
-`example_classifier_model.py 
<https://github.com/apache/airflow/blob/providers-common-ai/|version|/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_classifier_model.py>`__
-has both halves: a branch whose only classifier-specific line is ``model_id``, 
and a
+`example_decision_model.py 
<https://github.com/apache/airflow/blob/providers-common-ai/|version|/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_decision_model.py>`__
+has both halves: a branch with nothing specific to decision models but its 
connection, and a
 classification that escalates when the confidence is low.
 
 What it answers badly
 ---------------------
 
-Read `pydantic-ai's model page 
<https://pydantic.dev/docs/ai/models/typesafe/>`__ and
-`TypeSafe's own documentation <https://docs.typesafe.ai/>`__ before you
-trust a number from one of these models. Two of its failure modes matter more 
than the
-rest in a Dag:
+Read `pydantic-ai's decision model guide 
<https://pydantic.dev/docs/ai/models/decision/>`__
+and the documentation of the model you run (`TypeSafe's 
<https://docs.typesafe.ai/>`__ for
+Jev) before you trust a number from one of these models. Two failure modes 
documented for
+Jev matter more than the rest in a Dag; check them against any other model 
before relying on
+it not to share them:
 
 - **The text is treated as data, not as hostile.** An injected instruction, a 
misleading
   framing, or an argument for its own answer can move the result. A guard 
built on a
-  classifier model belongs alongside deterministic checks, not instead of them.
+  decision model belongs alongside deterministic checks, not instead of them.
 - **Option order is part of what the model sees.** Reordering a ``Literal``'s 
members can
   change the answer, so a threshold measured against one ordering is measured 
against
   that ordering only.
diff --git a/providers/common/ai/docs/examples.rst 
b/providers/common/ai/docs/examples.rst
index 592000c9938..0672b0d4b67 100644
--- a/providers/common/ai/docs/examples.rst
+++ b/providers/common/ai/docs/examples.rst
@@ -165,14 +165,14 @@ Reliability
      - What it shows
    * - :doc:`retry_policies`
      - Classifying task failures with an LLM into categories you define, then 
deriving
-       retry, fail, or delay from the category; and the same on a classifier 
model with a
+       retry, fail, or delay from the category; and the same on a decision 
model with a
        confidence bar. Source:
        `example_llm_retry_policy.py 
<https://github.com/apache/airflow/blob/providers-common-ai/|version|/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_llm_retry_policy.py>`__.
    * - :doc:`provider_fallback`
      - Failing over to another vendor inside one task attempt, and drilling 
the chain
        without waiting for an outage. Source:
        `example_llm_fallback.py 
<https://github.com/apache/airflow/blob/providers-common-ai/|version|/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_llm_fallback.py>`__.
-   * - :doc:`classifier_models`
+   * - :doc:`decision_models`
      - Routing a failure with a model that answers typed questions instead of 
writing
        text, and escalating when its confidence is low. Source:
-       `example_classifier_model.py 
<https://github.com/apache/airflow/blob/providers-common-ai/|version|/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_classifier_model.py>`__.
+       `example_decision_model.py 
<https://github.com/apache/airflow/blob/providers-common-ai/|version|/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_decision_model.py>`__.
diff --git a/providers/common/ai/docs/installation.rst 
b/providers/common/ai/docs/installation.rst
index 77e8b4e3e1f..a7227e724d6 100644
--- a/providers/common/ai/docs/installation.rst
+++ b/providers/common/ai/docs/installation.rst
@@ -40,8 +40,9 @@ The provider's extras split into a few groups:
 * **Model providers** (``openai``, ``anthropic``, ``google``, ``bedrock``, 
``typesafe``):
   pick the one matching your ``llm_conn_id`` connection 
(:doc:`model_providers` maps
   vendors to extras, prefixes and connection types). ``typesafe`` differs from 
the rest
-  in kind: it installs a classifier model that answers typed questions and 
cannot write
-  text (see :doc:`classifier_models`). The first four mirror the identically 
named
+  in kind: it installs a decision model that answers typed questions and 
cannot write
+  text (see :doc:`decision_models`). Other decision models, behind the System 
One API,
+  need no extra. The first four mirror the identically named
   ``pydantic-ai-slim`` optional dependency groups, and ``typesafe`` adds the 
``typesafe-sdk``
   the built-in adapter talks to; pydantic-ai supports more model providers
   than these, each under its own extra name, so check the
diff --git a/providers/common/ai/docs/model_providers.rst 
b/providers/common/ai/docs/model_providers.rst
index 7e31a70b648..fcfe587903a 100644
--- a/providers/common/ai/docs/model_providers.rst
+++ b/providers/common/ai/docs/model_providers.rst
@@ -92,11 +92,17 @@ type shown, and set the model name with that prefix.
      - ``pydanticai``
      - ``SNOWFLAKE_ACCOUNT`` and ``SNOWFLAKE_TOKEN`` in the worker 
environment; leave
        **Password** empty
-   * - TypeSafe Jev (classifier, does not write text)
+   * - TypeSafe Jev (decision model, does not write text)
      - ``typesafe:``
      - ``[typesafe]``
-     - ``pydanticai`` (:doc:`classifier_models`)
+     - ``pydanticai`` (:doc:`decision_models`)
      - API key in **Password**
+   * - Decision models behind the System One API: Ollama, Strands Decider, Kev 
and others
+       (do not write text)
+     - ``system-one:``
+     - None (``pydantic-ai-slim`` 2.53.0+)
+     - ``pydanticai`` with the server URL in **Host** (:doc:`decision_models`)
+     - API key in **Password**, if the server takes one
 
 ``[name]`` in the Install column is an extra of this provider, installed as
 ``pip install "apache-airflow-providers-common-ai[name]"``. The Groq, Mistral 
and Snowflake entries
@@ -141,7 +147,7 @@ Pages in this section
     AWS Bedrock <connections/pydantic_ai_bedrock>
     Google Vertex AI <connections/pydantic_ai_vertex>
     Self-hosted models <self_hosted_models>
-    Classifier models <classifier_models>
+    Decision models <decision_models>
     Provider fallback <provider_fallback>
     Connection reference <connections/pydantic_ai>
     Using the hook directly <hooks/pydantic_ai>
diff --git a/providers/common/ai/docs/operators/llm_branch.rst 
b/providers/common/ai/docs/operators/llm_branch.rst
index 6b07b1949a2..35b15f0f2b9 100644
--- a/providers/common/ai/docs/operators/llm_branch.rst
+++ b/providers/common/ai/docs/operators/llm_branch.rst
@@ -84,8 +84,7 @@ it, so the model can only answer with one of the task IDs.
 
 Descriptions explain the choices; they do not make the model more certain,
 and a text model's structured output carries no confidence to read. With a
-classifier model such as TypeSafe's, the descriptions become the criteria of
-its choice question, which is the text it weighs each option by.
+decision model, the descriptions become the criteria of its choice question, 
which is the text it weighs each option by.
 
 A pick is relative: the model chooses the best fit among the downstream tasks
 offered, not whether any of them fits. If "none of these" or "not enough to
@@ -176,8 +175,7 @@ Reviewing Uncertain Picks
     Experimental: this can change or be removed in a minor release of this 
provider.
     See :ref:`howto/stability`.
 
-A classifier model such as TypeSafe's returns a confidence with every pick,
-a number from 0 to 1 that summarizes how concentrated its probability
+A decision model returns a confidence with every pick, a number from 0 to 1 
that summarizes how concentrated its probability
 distribution was: near 1 when one branch stood out, low when two or more
 were close. It is not the probability that the pick is right. It is the
 model saying how clear-cut the question was, and it is the signal you gate on.
@@ -229,9 +227,9 @@ review it opens is the same one ``require_approval`` opens: 
``approval_timeout``
 ``on_approval_timeout``, ``allow_modifications``, ``approval_notifiers`` and
 ``approval_assigned_users`` all apply to it.
 
-TypeSafe's models need the provider's ``typesafe`` extra and a ``pydanticai``
-connection whose Model is ``typesafe:jev-1.13.0`` (or the ``model_id`` on the
-operator, as in the example). :doc:`../classifier_models` covers the setup.
+The example reads a ``pydanticai`` connection, ``decision_default``, whose 
Model is a decision
+model: ``typesafe:jev-1.13.0`` for TypeSafe's Jev, or ``system-one:<model>`` 
for a server
+answering the System One API. :doc:`../decision_models` covers the setup for 
each.
 
 The decision record
 ^^^^^^^^^^^^^^^^^^^
@@ -293,7 +291,7 @@ At execution time, the operator:
 3. Passes that type as ``output_type`` to ``pydantic-ai``, constraining the LLM
    to valid task IDs only.
 4. Reads the model's confidence for the pick from ``provider_details`` (a
-   classifier model reports one; a text model does not), pushes the
+   decision model reports one; a text model does not), pushes the
    ``decision`` XCom, and if the ``decision_policy`` or ``require_approval`` 
says
    so, pauses for human review (or fails, with ``on_uncertain="fail"``).
 5. Converts the LLM's structured output to task ID string(s) and calls
diff --git a/providers/common/ai/docs/redirects.txt 
b/providers/common/ai/docs/redirects.txt
index 295150c2603..a7c8c571403 100644
--- a/providers/common/ai/docs/redirects.txt
+++ b/providers/common/ai/docs/redirects.txt
@@ -22,3 +22,4 @@ sandbox.rst sandbox/index.rst
 hooks/index.rst concepts.rst
 end_to_end_pipelines.rst use_cases/index.rst
 guardrails.rst capabilities.rst
+classifier_models.rst decision_models.rst
diff --git a/providers/common/ai/docs/retry_policies.rst 
b/providers/common/ai/docs/retry_policies.rst
index ace62c46990..79625373783 100644
--- a/providers/common/ai/docs/retry_policies.rst
+++ b/providers/common/ai/docs/retry_policies.rst
@@ -42,8 +42,7 @@ fully model-driven, with the SDK's ``ExceptionRetryPolicy`` 
as the bottom rung:
        (``ClassifierRetryPolicy``)
      - The model names one of your ``categories``; the table says whether that
        category is retried, after how long, and how sure the model has to be.
-       A classifier model such as TypeSafe's Jev answers in a few hundred
-       milliseconds and reports its confidence; a text model can sit here too,
+       A decision model answers with its confidence; a text model can sit here 
too,
        without a bar.
      - Category descriptions and the confidence bar. No reasoning, no prose.
    * - **LLM**
@@ -174,7 +173,7 @@ failures belong there (the ``description`` the model 
reads), whether it is
 retried, after what ``delay``, and how sure the model has to be
 (``min_confidence``, covered below). Everything the model is told about a
 category, and everything the policy does with it, sits in that one entry, so
-the two cannot drift apart. This is also the policy a classifier model needs:
+the two cannot drift apart. This is also the policy a decision model needs:
 such a model refuses the free-text fields of ``ErrorClassification``, so an
 ``LLMRetryPolicy`` pointed at one fails every classification and falls back,
 with a log line saying to use ``ClassifierRetryPolicy``.
@@ -287,7 +286,7 @@ constructed, at Dag parse time, rather than on the first 
task failure.
 Confidence
 ----------
 
-A classifier model reports how sure it is of its answer. ``min_confidence`` is
+A decision model reports how sure it is of its answer. ``min_confidence`` is
 the bar that answer needs for the policy to act on it; under the bar the policy
 discards the answer and takes the same path it takes when the model call fails:
 ``fallback_rules`` if one matches, otherwise the task's own retry behaviour. It
@@ -296,7 +295,7 @@ does not substitute a delay of its own.
 Each category can carry its own bar. The stakes differ: a wrong ``transient``
 costs one more attempt, while a wrong ``permanent`` costs the task every retry 
it
 had left, so the category that ends the task deserves the higher bar.
-``jev_default`` is the classifier-model connection from 
:doc:`classifier_models`.
+``decision_default`` is the decision-model connection from 
:doc:`decision_models`.
 
 .. exampleinclude:: 
/../../ai/src/airflow/providers/common/ai/example_dags/example_llm_retry_policy.py
     :language: python
@@ -311,7 +310,7 @@ logs the confidence.
 confidence, and so does a response whose metadata was dropped along the way.
 With no bar configured that changes nothing. With a bar configured, every such
 answer is discarded and the fallback path decides, so swapping the connection
-from a classifier model to a text model does not silently switch off a control
+from a decision model to a text model does not silently switch off a control
 you set on purpose. To run a text model, remove the bar.
 
 The confidence is a statistic on the shape of the probability distribution the
@@ -323,7 +322,7 @@ answers start. A bar reduces wrong actions and does not 
eliminate them: a wrong
 pick can arrive with high confidence. Pin the model version
 (``typesafe:jev-1.13.0``, not ``jev-latest``): a bar tuned against one release
 is not guaranteed to mean the same thing after the next. See
-:doc:`classifier_models` for what these models answer well and badly.
+:doc:`decision_models` for what these models answer well and badly.
 
 Escalating to an LLM
 --------------------
@@ -507,7 +506,7 @@ When writing custom instructions:
   come back: a model that insists on one is re-prompted once by pydantic-ai and
   then gives up, which lands the task on ``fallback_rules`` or on its own retry
   behaviour, having billed two calls.
-- A classifier model sends ``instructions`` as the question it scores the
+- A decision model sends ``instructions`` as the question it scores the
   exception text against, not as rules it follows step by step, so a long 
rubric
   buys less there than a better description on each category does.
 
diff --git a/providers/common/ai/docs/self_hosted_models.rst 
b/providers/common/ai/docs/self_hosted_models.rst
index ef1244f78f2..f454d0a2a2b 100644
--- a/providers/common/ai/docs/self_hosted_models.rst
+++ b/providers/common/ai/docs/self_hosted_models.rst
@@ -374,6 +374,9 @@ Where to go next
 -------------------
 
 - :ref:`howto/connection:pydanticai` -- the full connection field reference.
+- :doc:`decision_models` -- a self-hosted decision model, such as Strands 
Decider or one
+  Ollama runs, takes ``system-one:<model>`` with the server URL in ``host`` 
rather than
+  the ``openai:`` or ``ollama:`` prefixes on this page.
 - :doc:`retry_policies` -- the "Local LLM support" section covers pointing
   ``LLMRetryPolicy`` at a self-hosted endpoint.
 - :doc:`examples` -- more runnable Dags against the ``pydanticai``
diff --git a/providers/common/ai/docs/stability.rst 
b/providers/common/ai/docs/stability.rst
index 100ba376555..2101c651253 100644
--- a/providers/common/ai/docs/stability.rst
+++ b/providers/common/ai/docs/stability.rst
@@ -111,7 +111,7 @@ Everything this provider ships that is not in the table 
above is experimental.
    * - Feature
      - Why it is experimental
    * - 
:class:`~airflow.providers.common.ai.policies.retry.ClassifierRetryPolicy`
-       (:doc:`classifier_models`)
+       (:doc:`decision_models`)
      - The confidence threshold, the fallback order and the behaviour when the 
classifier
        is unavailable are still settling.
    * - :class:`~airflow.providers.common.ai.policies.decision.DecisionPolicy`, 
and
diff --git a/providers/common/ai/docs/use_cases/index.rst 
b/providers/common/ai/docs/use_cases/index.rst
index 027d79c33e4..7793e4e4c5d 100644
--- a/providers/common/ai/docs/use_cases/index.rst
+++ b/providers/common/ai/docs/use_cases/index.rst
@@ -30,7 +30,8 @@ the shape most of the others build on: structured output plus 
dynamic task mappi
 
 Every Dag needs the provider installed with the extra for your model vendor 
and a
 ``pydanticai`` connection named ``pydanticai_default``; :doc:`../quickstart` 
covers both.
-Each page's "Run it" lists only what that Dag adds.
+Each page's "Run it" lists only what that Dag adds. 
:doc:`route_pipeline_failures` is the
+exception: it reads a decision-model connection, ``decision_default``, instead.
 
 .. list-table::
    :header-rows: 1
diff --git a/providers/common/ai/docs/use_cases/route_pipeline_failures.rst 
b/providers/common/ai/docs/use_cases/route_pipeline_failures.rst
index 046961a8cbf..10ca73f304a 100644
--- a/providers/common/ai/docs/use_cases/route_pipeline_failures.rst
+++ b/providers/common/ai/docs/use_cases/route_pipeline_failures.rst
@@ -28,23 +28,18 @@ What this demonstrates
 ----------------------
 
 * :doc:`../operators/llm_branch` -- ``LLMBranchOperator`` picks one downstream 
task id.
-* :doc:`../classifier_models` -- a classifier model, so the gate has a 
confidence to read.
+* :doc:`../decision_models` -- a decision model, so the gate has a confidence 
to read.
 * :doc:`../approval_gates` -- ``on_uncertain="review"`` routes low-confidence 
picks to a
   human, with ``approval_timeout`` bounding the wait.
 
 Run it
 ------
 
-1. Install the provider with the classifier extra:
+1. Create the ``decision_default`` connection for the decision model you run, 
TypeSafe Jev or
+   a server answering the System One API (:doc:`../decision_models` covers 
both). Both Dags
+   below read it.
 
-   .. code-block:: bash
-
-       pip install "apache-airflow-providers-common-ai[typesafe]"
-
-2. Set ``{"model": "typesafe:jev-1.13.0"}`` in the ``pydanticai_default`` 
extra. The second
-   Dag below reads the same kind of connection as ``jev_default``.
-
-3. Trigger the Dag:
+2. Trigger the Dag:
 
    .. code-block:: bash
 
@@ -67,13 +62,13 @@ Classify-then-act variant
 -------------------------
 
 When the action depends on the score itself, classify in one task and act in 
the next
-(:doc:`../classifier_models` explains reading the score). It uses the 
``jev_default``
-connection; run it as ``airflow dags test 
example_classifier_model_confidence``:
+(:doc:`../decision_models` explains reading the score). It uses the same 
``decision_default``
+connection; run it as ``airflow dags test example_decision_model_confidence``:
 
-.. exampleinclude:: 
/../../ai/src/airflow/providers/common/ai/example_dags/example_classifier_model.py
+.. exampleinclude:: 
/../../ai/src/airflow/providers/common/ai/example_dags/example_decision_model.py
     :language: python
-    :start-after: [START howto_classifier_model_confidence]
-    :end-before: [END howto_classifier_model_confidence]
+    :start-after: [START howto_decision_model_confidence]
+    :end-before: [END howto_decision_model_confidence]
 
 Adapting it
 -----------
diff --git 
a/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_classifier_model.py
 
b/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_decision_model.py
similarity index 60%
rename from 
providers/common/ai/src/airflow/providers/common/ai/example_dags/example_classifier_model.py
rename to 
providers/common/ai/src/airflow/providers/common/ai/example_dags/example_decision_model.py
index b56124f50fc..179d7bf3066 100644
--- 
a/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_classifier_model.py
+++ 
b/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_decision_model.py
@@ -15,35 +15,41 @@
 # specific language governing permissions and limitations
 # under the License.
 """
-Example DAG using a classifier model to route an incident.
+Example DAG using a decision model to route an incident.
 
-A classifier model answers typed questions and cannot write text, so it suits 
a branch
+A decision model answers typed questions and cannot write text, so it suits a 
branch
 (pick one of these task ids) and a classification (pick one of these labels), 
and nothing
-that produces prose. See the "Classifier models" guide in this provider's 
documentation.
+that produces prose. See the "Decision models" guide in this provider's 
documentation.
 
-Prerequisites:
-  - ``pip install 'apache-airflow-providers-common-ai[typesafe]'``
-  - Connection ``jev_default`` with ``conn_type='pydanticai'``, 
``password=<API key>``,
-    ``extra='{"model": "typesafe:jev-1.13.0"}'``
+Prerequisites: ``pydantic-ai-slim`` 2.46.0 or later (2.53.0 for 
``system-one:``), and a
+connection ``decision_default`` with ``conn_type='pydanticai'`` whose model is 
a decision model,
+either:
+
+- TypeSafe Jev: ``password=<API key>``, ``extra='{"model": 
"typesafe:jev-1.13.0"}'``, and
+  ``pip install 'apache-airflow-providers-common-ai[typesafe]'``, or
+- a server answering the System One API: ``host=<server URL>``, 
``password=<API key, if any>``,
+  ``extra='{"model": "system-one:<model name>"}'``
 """
 
 from __future__ import annotations
 
-from typing import Literal
+from enum import Enum
+
+from pydantic_ai import UseEnumMemberDocstrings
 
 from airflow.providers.common.ai.hooks.pydantic_ai import PydanticAIHook
 from airflow.providers.common.ai.operators.llm_branch import LLMBranchOperator
 from airflow.providers.common.compat.sdk import dag, task
 
 # Clear-cut: the error names the missing permission, so one remediation fits 
and the others
-# do not. Branch options a classifier model can separate are the ones to give 
it.
+# do not. Branch options a decision model can separate are the ones to give it.
 PERMISSION_FAILURE = (
     "Task export_report failed: botocore.exceptions.ClientError: An error 
occurred "
     "(AccessDenied) when calling the PutObject operation: the task role "
     "airflow-worker lacks s3:PutObject on arn:aws:s3:::reports-prod/daily/."
 )
 
-# Answerable but not decisively: a classifier model lands on ``resource`` here 
with roughly
+# Answerable but not decisively: a decision model lands on ``resource`` here 
with roughly
 # two thirds of the probability, which is enough to be worth recording and not 
enough to act
 # on unattended. That is the case the confidence gate below exists for.
 UNDERDETERMINED_INCIDENT = (
@@ -59,16 +65,22 @@ ACT_ABOVE = 0.8
 REVIEW_ABOVE = 0.5
 
 
-# [START howto_classifier_model_branch]
-@dag(tags=["example", "classifier"])
-def example_classifier_model_branch():
-    """Route the incident. The model id is the only thing that makes this a 
classifier."""
+# [START howto_decision_model_branch]
+@dag(tags=["example", "decision"])
+def example_decision_model_branch():
+    """Route the incident. The connection's model is the only thing that makes 
this a decision model."""
     route = LLMBranchOperator(
         task_id="route_failure",
         prompt=PERMISSION_FAILURE,
-        llm_conn_id="jev_default",
-        model_id="typesafe:jev-1.13.0",
+        llm_conn_id="decision_default",
         system_prompt="Pick the remediation that addresses the cause, not the 
symptom.",
+        # Each description is what the model weighs that branch by. Some 
System One servers
+        # refuse an option without one, so describe every branch.
+        branches={
+            "grant_bucket_write": "The task's role lacks a permission it needs 
on the bucket.",
+            "restore_deleted_bucket": "The destination bucket or prefix no 
longer exists.",
+            "wait_and_retry": "A transient fault: throttling, a timeout, a 
brief outage.",
+        },
     )
 
     @task
@@ -86,14 +98,29 @@ def example_classifier_model_branch():
     route >> [grant_bucket_write(), restore_deleted_bucket(), wait_and_retry()]
 
 
-# [END howto_classifier_model_branch]
+# [END howto_decision_model_branch]
+
+example_decision_model_branch()
+
+
+# [START howto_decision_model_confidence]
+class FailureCause(UseEnumMemberDocstrings, str, Enum):
+    """Why an Airflow task failed."""
+
+    # Each docstring is what the model weighs that option by. Some System One 
servers
+    # refuse an option without one, so describe every option.
+    transient = "transient"
+    """A fault that clears by itself: a timeout, throttling, a dropped 
connection."""
+
+    resource = "resource"
+    """A dependency is down or unreachable and needs fixing before a retry can 
work."""
 
-example_classifier_model_branch()
+    permanent = "permanent"
+    """A bug or bad input that fails the same way however often it is 
retried."""
 
 
-# [START howto_classifier_model_confidence]
-@dag(tags=["example", "classifier"])
-def example_classifier_model_confidence():
+@dag(tags=["example", "decision"])
+def example_decision_model_confidence():
     """Classify a failure and escalate when the model says it does not know.
 
     ``LLMBranchOperator`` can gate its pick on one confidence bar through 
``decision_policy``.
@@ -103,8 +130,8 @@ def example_classifier_model_confidence():
 
     @task
     def classify(log_line: str) -> dict:
-        agent = PydanticAIHook(llm_conn_id="jev_default").create_agent(
-            output_type=Literal["transient", "resource", "permanent"],
+        agent = PydanticAIHook(llm_conn_id="decision_default").create_agent(
+            output_type=FailureCause,
             instructions="Classify why this Airflow task failed.",
         )
         result = agent.run_sync(log_line)
@@ -115,7 +142,7 @@ def example_classifier_model_confidence():
         # same confidence in their ``decision`` XCom.
         details = result.response.provider_details or {}
         return {
-            "category": result.output,
+            "category": result.output.value,
             "confidence": (details.get("confidence") or {}).get("response"),
         }
 
@@ -131,6 +158,6 @@ def example_classifier_model_confidence():
     act(classify(UNDERDETERMINED_INCIDENT))
 
 
-# [END howto_classifier_model_confidence]
+# [END howto_decision_model_confidence]
 
-example_classifier_model_confidence()
+example_decision_model_confidence()
diff --git 
a/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_llm_branch.py
 
b/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_llm_branch.py
index 0a17f9c43a1..d210395acc9 100644
--- 
a/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_llm_branch.py
+++ 
b/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_llm_branch.py
@@ -101,7 +101,7 @@ example_llm_branch_descriptions()
 # [START howto_operator_llm_branch_decision_policy]
 @dag(tags=["example"])
 def example_llm_branch_decision_policy():
-    # A classifier model reports how sure it is of each pick; a text model 
does not, and
+    # A decision model reports how sure it is of each pick; a text model does 
not, and
     # with a min_confidence set every pick would count as uncertain and go to 
review.
     route = LLMBranchOperator(
         task_id="triage_failure",
@@ -109,8 +109,7 @@ def example_llm_branch_decision_policy():
             "Task load_orders failed: psycopg2.OperationalError: could not 
connect to server: "
             "Connection timed out. Is the server running on host db.internal 
(10.0.4.12)?"
         ),
-        llm_conn_id="pydanticai_default",
-        model_id="typesafe:jev-1.13.0",
+        llm_conn_id="decision_default",
         system_prompt="Pick the remediation that addresses the cause of the 
failure.",
         branches={
             "rerun": "The failure looks transient: a timeout, a dropped 
connection, a rate limit.",
diff --git 
a/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_llm_retry_policy.py
 
b/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_llm_retry_policy.py
index 869fee2e549..a2dd6a52444 100644
--- 
a/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_llm_retry_policy.py
+++ 
b/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_llm_retry_policy.py
@@ -19,16 +19,16 @@ Example DAG demonstrating LLM-powered retry policies.
 
 ``llm_policy`` is the plain form: a text model classifies the failure and 
decides whether to
 retry and how long to wait, guided by its instructions. ``patient_policy`` and 
the
-classifier-model policies are ``ClassifierRetryPolicy``: the model only names 
the kind of failure
+decision-model policies are ``ClassifierRetryPolicy``: the model only names 
the kind of failure
 and the policy's table decides the action, the delay and how sure the model 
has to be.
 
 Prerequisites:
   - Connection ``pydanticai_default`` with ``conn_type='pydanticai'``,
     ``password=<API key>``, ``extra='{"model": 
"anthropic:claude-haiku-4-5-20251001"}'``
   - ``pip install apache-airflow-providers-common-ai[anthropic]``
-  - For the classifier-model Dag: connection ``jev_default`` with
-    ``extra='{"model": "typesafe:jev-1.13.0"}'`` and
-    ``pip install 'apache-airflow-providers-common-ai[typesafe]'``
+  - For the decision-model Dag: a connection ``decision_default`` whose model 
is a decision
+    model, as set up in the "Decision models" guide. Its confidence bars were 
calibrated on
+    ``typesafe:jev-1.13.0``; measure them again for another model.
 """
 
 from __future__ import annotations
@@ -96,15 +96,15 @@ try:
     example_llm_retry_policy()
 
     # [START howto_retry_policy_classifier]
-    # A classifier model answers the same question in a few hundred 
milliseconds and reports
-    # how sure it is. The categories are this pipeline's own; their 
descriptions are what the
+    # A decision model answers the same question with how sure it is, and 
jev-1.13.0 does it
+    # in a few hundred milliseconds. The categories are this pipeline's own; 
their descriptions are what the
     # model reads. ``permanent`` ends the task and costs it every retry it had 
left, so it
     # demands more certainty than the rest. Under a bar the answer is 
discarded and the
     # fallback rules, then the task's own retry settings, decide instead. The 
bars come from
     # a calibration run on jev-1.13.0: correct picks landed at 0.89 and above, 
wrong ones
     # at 0.47 to 0.69, with one wrong ``permanent`` at 0.90 that no sensible 
bar catches.
     snowflake_policy = ClassifierRetryPolicy(
-        llm_conn_id="jev_default",
+        llm_conn_id="decision_default",
         min_confidence=0.8,
         categories={
             "queued": ErrorCategory(
@@ -134,11 +134,11 @@ try:
     # [END howto_retry_policy_classifier]
 
     # [START howto_retry_policy_escalation]
-    # The three layers chained. The classifier answers the clear cases in a 
few hundred
-    # milliseconds. When it is unsure or unreachable, a text model reasons 
about the failure
-    # and decides retry and delay itself. When that model is unreachable too, 
the rules decide.
+    # The three layers chained. The decision model answers the clear cases 
cheaply. When it is
+    # unsure or unreachable, a text model reasons about the failure and 
decides retry and delay
+    # itself. When that model is unreachable too, the rules decide.
     escalating_policy = ClassifierRetryPolicy(
-        llm_conn_id="jev_default",
+        llm_conn_id="decision_default",
         min_confidence=0.8,
         categories=snowflake_policy.categories,
         fallback_policy=LLMRetryPolicy(llm_conn_id="pydanticai_default", 
timeout=30.0),
diff --git 
a/providers/common/ai/src/airflow/providers/common/ai/operators/llm.py 
b/providers/common/ai/src/airflow/providers/common/ai/operators/llm.py
index f2ce27a315a..3c11b3bd1a8 100644
--- a/providers/common/ai/src/airflow/providers/common/ai/operators/llm.py
+++ b/providers/common/ai/src/airflow/providers/common/ai/operators/llm.py
@@ -158,10 +158,10 @@ class LLMOperator(CancellableAgentRunMixin, BaseOperator, 
LLMApprovalMixin):
         saying how confident the model has to be for the operator to return 
its answer by
         itself (``min_confidence``) and what happens otherwise 
(``on_uncertain``: ``"review"``
         or ``"fail"``). Confidence comes from models that report one per 
output field, such as
-        a classifier model (TypeSafe's), in ``provider_details``. A structured 
output is judged
-        by its least confident field among the fields that reported one; a 
field whose type
-        reports none (a bounded float, where the probability is the answer) is 
not gated, and
-        the record's ``confidence`` map shows which fields were compared. When 
no field reports
+        a decision model, in ``provider_details``. A structured output is 
judged by its least
+        confident field among the fields that reported one; a field whose type 
reports none (a
+        bounded float, where the probability is the answer) is not gated, and 
the record's
+        ``confidence`` map shows which fields were compared. When no field 
reports
         any confidence, as with a text model, the output counts as uncertain, 
so swapping the
         connection does not silently switch off a control the author set. 
Independent of
         ``require_approval``, which always asks. ``on_uncertain="review"`` 
needs Airflow 3.1+,
diff --git 
a/providers/common/ai/src/airflow/providers/common/ai/operators/llm_branch.py 
b/providers/common/ai/src/airflow/providers/common/ai/operators/llm_branch.py
index 02f603ba3af..2942b1601d4 100644
--- 
a/providers/common/ai/src/airflow/providers/common/ai/operators/llm_branch.py
+++ 
b/providers/common/ai/src/airflow/providers/common/ai/operators/llm_branch.py
@@ -98,19 +98,19 @@ class LLMBranchOperator(LLMOperator, BranchMixIn):
         output schema next to the option, so the model reads "here is an 
option, here is
         what it means" rather than guessing from the task ID; 
``min_confidence`` on an
         option, which is experimental, is a bar for that branch alone. A 
downstream task
-        without an entry is
-        presented by its ID alone and takes the policy's bar. A key that is 
not a
-        downstream task ID fails the task before the model is called. 
Descriptions
-        support Jinja templating.
+        without an entry is presented by its ID alone and takes the policy's 
bar; some
+        decision-model servers refuse an option without a description, so 
describe every
+        branch when the connection is a decision model. A key that is not a 
downstream task
+        ID fails the task before the model is called. Descriptions support 
Jinja templating.
     :param allow_multiple_branches: When ``False`` (default) the LLM returns a
         single task ID. When ``True`` the LLM may return one or more task IDs.
     :param decision_policy: Experimental. A
         :class:`~airflow.providers.common.ai.policies.decision.DecisionPolicy`:
         the confidence a pick needs for the operator to branch on it without a 
person
         (``min_confidence``) and what happens under it (``on_uncertain``: 
``"review"`` or
-        ``"fail"``). Confidence comes from models that report one, such as a 
classifier
-        model (TypeSafe's); a text model reports none, which counts as 
uncertain, so
-        swapping the connection does not silently switch off a control the 
author set.
+        ``"fail"``). Confidence comes from models that report one, such as a 
decision model;
+        a text model reports none, which counts as uncertain, so swapping the 
connection does
+        not silently switch off a control the author set.
         A branch's own ``min_confidence`` overrides the policy's for that 
pick; with
         ``allow_multiple_branches`` the strictest bar among the picked 
branches applies.
         ``require_approval=True`` still sends every pick to a person 
regardless.
diff --git 
a/providers/common/ai/src/airflow/providers/common/ai/policies/decision.py 
b/providers/common/ai/src/airflow/providers/common/ai/policies/decision.py
index 4e60b2bc2ad..68494c82238 100644
--- a/providers/common/ai/src/airflow/providers/common/ai/policies/decision.py
+++ b/providers/common/ai/src/airflow/providers/common/ai/policies/decision.py
@@ -56,8 +56,8 @@ class DecisionPolicy:
 
     :param min_confidence: The confidence, from 0 to 1, the answer needs for 
the operator to act
         without a person. ``None`` (default) is no gate: the operator behaves 
as it always has.
-        Confidence comes from models that report one, such as a classifier 
model (TypeSafe's),
-        in ``provider_details``. A model that reports none counts as under the 
bar, so swapping
+        Confidence comes from models that report one, such as a decision 
model, in
+        ``provider_details``. A model that reports none counts as under the 
bar, so swapping
         the connection to a text model does not silently switch off a control 
the author set.
     :param on_uncertain: What happens under the bar. ``"review"`` (default) 
sends the answer to
         human review through the same approval flow as ``require_approval``, 
which needs Airflow
diff --git 
a/providers/common/ai/src/airflow/providers/common/ai/policies/retry.py 
b/providers/common/ai/src/airflow/providers/common/ai/policies/retry.py
index ad21d52b569..6440dcbaec6 100644
--- a/providers/common/ai/src/airflow/providers/common/ai/policies/retry.py
+++ b/providers/common/ai/src/airflow/providers/common/ai/policies/retry.py
@@ -23,8 +23,7 @@ Model-backed retry policies, one per layer of a ladder from 
hardcoded to reasoni
 * **Classifier.** :class:`ClassifierRetryPolicy`: the model names one of the 
author's
   ``categories`` and the :class:`ErrorCategory` table decides whether that 
category is
   retried, after how long, and how sure the model has to be. Tuned through 
descriptions and
-  a confidence bar, not through reasoning. A classifier model such as 
TypeSafe's Jev runs
-  here; a text model can too.
+  a confidence bar, not through reasoning. A decision model runs here; a text 
model can too.
 * **LLM.** :class:`LLMRetryPolicy`: a text model classifies the failure, 
decides whether to
   retry and how long to wait from ``instructions``, and explains itself.
 
@@ -347,9 +346,8 @@ class LLMRetryPolicy(_ModelRetryPolicy):
     Ollama, etc.) for error classification with structured output. The model
     returns an :class:`ErrorClassification`: which category the error is, 
whether
     to retry, how long to wait, and why, all steered by ``instructions``. This 
is
-    the reasoning layer; for a cheap typed decision from a classifier model 
such as
-    TypeSafe's Jev, use :class:`ClassifierRetryPolicy`, which can name this 
policy as
-    its ``fallback_policy``.
+    the reasoning layer; for a cheap typed decision from a decision model, use
+    :class:`ClassifierRetryPolicy`, which can name this policy as its 
``fallback_policy``.
 
     When the LLM call itself fails, the policy falls back to ``fallback_rules``
     (if provided) or returns DEFAULT to use the task's standard retry logic.
@@ -398,11 +396,11 @@ class LLMRetryPolicy(_ModelRetryPolicy):
             result = self._run(agent, exception, try_number, max_tries)
         except Exception as exc:
             if "not supported by this model" in str(exc):
-                # A classifier model refuses ErrorClassification's free-text 
fields client-side. A text
+                # A decision model refuses ErrorClassification's free-text 
fields client-side. A text
                 # model whose profile lacks structured output raises the same 
words, so this is a hint.
                 log.error(
-                    "This model cannot answer ErrorClassification. If it is a 
classifier model such as "
-                    "TypeSafe's Jev, use ClassifierRetryPolicy, which asks it 
a typed question."
+                    "This model cannot answer ErrorClassification. If it is a 
decision model, "
+                    "use ClassifierRetryPolicy, which asks it a typed 
question."
                 )
             raise
         classification = result.output
@@ -441,8 +439,7 @@ class ClassifierRetryPolicy(_ModelRetryPolicy):
     The model's only job is to pick one of ``categories``; it reads each one's 
description
     from the output schema. Whether that category is retried, after how long, 
and how sure
     the model has to be all come from the :class:`ErrorCategory` in the worker 
process.
-    That is the shape a classifier model such as TypeSafe's Jev answers, in a 
few hundred
-    milliseconds and with a confidence; a text model answers it too.
+    That is the shape a decision model answers, with a confidence; a text 
model answers it too.
 
     When the model call fails, or the answer is under its confidence bar, the 
policy
     consults ``fallback_policy`` if set, then ``fallback_rules``, then returns 
DEFAULT to use
@@ -467,7 +464,7 @@ class ClassifierRetryPolicy(_ModelRetryPolicy):
         outside them is rejected before the policy acts on it.
     :param min_confidence: The confidence, from 0 to 1, the model's answer 
needs for the
         policy to act on it. ``None`` (default) is no bar: the answer is acted 
on whatever
-        the confidence. Confidence comes from models that report one, such as 
a classifier
+        the confidence. Confidence comes from models that report one, such as 
a decision
         model, in ``provider_details``. Under the bar, or when a bar is set 
and the model
         reported no confidence, the answer is discarded and 
``fallback_policy``, then
         ``fallback_rules``, then the task's own retry behaviour apply, so 
swapping the
diff --git 
a/providers/common/ai/src/airflow/providers/common/ai/utils/decision.py 
b/providers/common/ai/src/airflow/providers/common/ai/utils/decision.py
index 6d7a94d49a3..32193d1c2fd 100644
--- a/providers/common/ai/src/airflow/providers/common/ai/utils/decision.py
+++ b/providers/common/ai/src/airflow/providers/common/ai/utils/decision.py
@@ -17,9 +17,9 @@
 """
 Confidence gating and the decision record shared by the LLM operators.
 
-A classifier model such as TypeSafe's reports a confidence per output field in
-``ModelResponse.provider_details`` (``{"confidence": {field: float}, 
"probabilities":
-{field: {option: float}}}``). A text model reports nothing there. These 
helpers read
+A decision model reports a confidence per output field in 
``ModelResponse.provider_details``
+(``{"confidence": {field: float}, "probabilities": {field: {option: 
float}}}``).
+A text model reports nothing there. These helpers read
 that, decide whether an answer should go to a person before the operator acts 
on it,
 and shape the ``decision`` XCom record that says what was proposed, what was 
done, and
 why.
@@ -71,7 +71,7 @@ __all__ = [
 DECISION_XCOM_KEY = "decision"
 
 BARE_OUTPUT_FIELD = "response"
-"""The field name a classifier model reports a bare (non-object) output type's 
confidence under."""
+"""The field name a decision model reports a bare (non-object) output type's 
confidence under."""
 
 ReviewReason = Literal["require_approval", "below_threshold", 
"missing_confidence"]
 DecidedBy = Literal["model", "human", "timeout_default", "policy", "timeout"]
@@ -236,7 +236,7 @@ def check_uncertain_action(
         raise LowConfidenceError(
             f"min_confidence={threshold:.2f} is set for {what} but the model 
reported no confidence for "
             "its answer, and on_uncertain='fail'. A text model reports none; a 
bounded float field "
-            "reports none because the probability is the answer. Use a 
classifier model, remove the "
+            "reports none because the probability is the answer. Use a 
decision model, remove the "
             "bar, or set on_uncertain='review'."
         )
     raise LowConfidenceError(
@@ -311,7 +311,7 @@ def described_choices(name: str, options: Mapping[str, str 
| None]) -> type[Any]
 
     With a description on any option the schema renders as ``anyOf`` of 
``{const, description}``
     instead of a bare ``enum`` list. That is the one JSON Schema shape that 
carries a description
-    per value, and it is what both a text model's tool schema and 
pydantic-ai's TypeSafe adapter
+    per value, and it is what both a text model's tool schema and 
pydantic-ai's decision models
     read an option's meaning from. On pydantic-ai 2.46+ the type is its 
``Choices``; before that,
     an ``Enum`` whose schema hook emits the same shape. Either way the model 
has to answer with one
     of the keys, in the order given, and :func:`picked_key` returns that key 
whichever type answered.
diff --git 
a/providers/common/ai/tests/unit/common/ai/operators/test_llm_branch.py 
b/providers/common/ai/tests/unit/common/ai/operators/test_llm_branch.py
index 0ea1dd6cb30..5ecb7d26abe 100644
--- a/providers/common/ai/tests/unit/common/ai/operators/test_llm_branch.py
+++ b/providers/common/ai/tests/unit/common/ai/operators/test_llm_branch.py
@@ -211,7 +211,7 @@ class TestLLMBranchOperator:
         """The option order the model sees is sorted, not whatever order 
downstream_task_ids iterates in.
 
         ``downstream_task_ids`` is a set, so its iteration order depends on 
string hashing and
-        differs between worker processes. Option order is part of the question 
for a classifier
+        differs between worker processes. Option order is part of the question 
for a decision
         model, so it has to be the same on every worker. A reverse-sorted list 
stands in for an
         unlucky set order; without ``sorted()`` the enum comes out reversed 
and this fails.
         """
diff --git a/providers/common/ai/tests/unit/common/ai/policies/test_retry.py 
b/providers/common/ai/tests/unit/common/ai/policies/test_retry.py
index 5168b6b1c4f..30ca9d2a679 100644
--- a/providers/common/ai/tests/unit/common/ai/policies/test_retry.py
+++ b/providers/common/ai/tests/unit/common/ai/policies/test_retry.py
@@ -1220,7 +1220,7 @@ class TestLLMRetryPolicy:
 
     @patch(HOOK, autospec=True)
     def test_classifier_refusal_logs_the_categories_hint(self, mock_hook_cls, 
caplog):
-        """A classifier model refuses ErrorClassification's text fields; the 
log says what to do."""
+        """A decision model refuses ErrorClassification's text fields; the log 
says what to do."""
         agent = _install(mock_hook_cls, MagicMock(spec=Agent))
         agent.run_sync.side_effect = RuntimeError("Output field 'reasoning' is 
not supported by this model")
         policy = LLMRetryPolicy(llm_conn_id="test", 
model_id="typesafe:jev-1.13.0")

Reply via email to