This is an automated email from the ASF dual-hosted git repository.
kaxil pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/airflow.git
The following commit(s) were added to refs/heads/main by this push:
new 758ebbd35f4 Document decision models served over the System One API in
common.ai (#74266)
758ebbd35f4 is described below
commit 758ebbd35f4ad963d29058a326c719626d9f6000
Author: Kaxil Naik <[email protected]>
AuthorDate: Mon Oct 5 21:00:57 2026 +0100
Document decision models served over the System One API in common.ai
(#74266)
pydantic-ai 2.53.0 added SystemOneModel, which runs any decision model
behind the POST /v1/systemone API Jev uses (Ollama, Strands Decider, Kev
and others). PydanticAIHook already builds it from a pydanticai connection
through its generic provider path, so this is docs and examples only.
Rename the "Classifier models" page to "Decision models" with a redirect,
cover both the typesafe: and system-one: backends in setup, and make the
decision-model examples read a decision_default connection with every
option described, so they run unchanged on either backend.
Co-authored-by: Claude <[email protected]>
---
docs/spelling_wordlist.txt | 1 +
providers/common/ai/docs/approval_gates.rst | 4 +-
.../{classifier_models.rst => decision_models.rst} | 158 +++++++++++++++------
providers/common/ai/docs/examples.rst | 6 +-
providers/common/ai/docs/installation.rst | 5 +-
providers/common/ai/docs/model_providers.rst | 12 +-
providers/common/ai/docs/operators/llm_branch.rst | 14 +-
providers/common/ai/docs/redirects.txt | 1 +
providers/common/ai/docs/retry_policies.rst | 15 +-
providers/common/ai/docs/self_hosted_models.rst | 3 +
providers/common/ai/docs/stability.rst | 2 +-
providers/common/ai/docs/use_cases/index.rst | 3 +-
.../ai/docs/use_cases/route_pipeline_failures.rst | 25 ++--
...assifier_model.py => example_decision_model.py} | 79 +++++++----
.../common/ai/example_dags/example_llm_branch.py | 5 +-
.../ai/example_dags/example_llm_retry_policy.py | 22 +--
.../airflow/providers/common/ai/operators/llm.py | 8 +-
.../providers/common/ai/operators/llm_branch.py | 14 +-
.../providers/common/ai/policies/decision.py | 4 +-
.../airflow/providers/common/ai/policies/retry.py | 19 ++-
.../airflow/providers/common/ai/utils/decision.py | 12 +-
.../unit/common/ai/operators/test_llm_branch.py | 2 +-
.../ai/tests/unit/common/ai/policies/test_retry.py | 2 +-
23 files changed, 260 insertions(+), 156 deletions(-)
diff --git a/docs/spelling_wordlist.txt b/docs/spelling_wordlist.txt
index 6db426432a6..325b68d6c90 100644
--- a/docs/spelling_wordlist.txt
+++ b/docs/spelling_wordlist.txt
@@ -972,6 +972,7 @@ kerberized
Kerberos
kerberos
KerberosClient
+Kev
keycloak
Keyfile
keyfile
diff --git a/providers/common/ai/docs/approval_gates.rst
b/providers/common/ai/docs/approval_gates.rst
index 39b16ea09b5..c40eee8e2e9 100644
--- a/providers/common/ai/docs/approval_gates.rst
+++ b/providers/common/ai/docs/approval_gates.rst
@@ -90,8 +90,8 @@ Reviewing uncertain output
Experimental: this can change or be removed in a minor release of this
provider.
See :ref:`howto/stability`.
-A classifier model such as TypeSafe's reports a confidence for every field
-of a structured output, in ``provider_details`` on the model response. It is a
+A decision model reports a confidence for every field of a structured output,
in
+``provider_details`` on the model response. It is a
summary of how concentrated the model's probability distribution was, not the
probability that the field is right.
``decision_policy=DecisionPolicy(min_confidence=0.7)``
(import ``DecisionPolicy`` from ``airflow.providers.common.ai.operators.llm``)
diff --git a/providers/common/ai/docs/classifier_models.rst
b/providers/common/ai/docs/decision_models.rst
similarity index 52%
rename from providers/common/ai/docs/classifier_models.rst
rename to providers/common/ai/docs/decision_models.rst
index 2f003806e56..3f32da16aa2 100644
--- a/providers/common/ai/docs/classifier_models.rst
+++ b/providers/common/ai/docs/decision_models.rst
@@ -15,53 +15,110 @@
specific language governing permissions and limitations
under the License.
-Classifier models
-=================
+Decision models
+===============
.. note::
Experimental: this can change or be removed in a minor release of this
provider.
See :ref:`howto/stability`.
-Every other model in this provider writes text. A classifier model does not:
you give it
+Every other model in this provider writes text. A decision model does not: you
give it
some text and a typed question, and it answers with a value from a set you
named in
advance, plus a confidence. Ask it for a string and the request is refused
before it
leaves your process.
-`TypeSafe <https://typesafe.ai>`__'s Jev is the one pydantic-ai supports, as
the
-``typesafe:`` provider. Nothing in this provider is specific to it -- it
arrives through the same
+pydantic-ai reaches decision models through two model prefixes. Nothing in
this provider is
+specific to either: both arrive through the same
:class:`~airflow.providers.common.ai.hooks.pydantic_ai.PydanticAIHook` as
every other
model, so a model id is the whole integration.
+.. list-table::
+ :header-rows: 1
+ :widths: 18 32 50
+
+ * - Prefix
+ - Needs
+ - Runs
+ * - ``typesafe:``
+ - ``pydantic-ai-slim`` 2.45.0 or later (2.46.0 for the examples on this
page, which
+ describe options with ``UseEnumMemberDocstrings``) and the provider's
``typesafe`` extra
+ - `TypeSafe <https://typesafe.ai>`__'s hosted Jev.
+ * - ``system-one:``
+ - ``pydantic-ai-slim`` 2.53.0 or later, no extra
+ - Any server that answers the same ``POST /v1/systemone`` API as Jev,
hosted or your
+ own. `Ollama <https://docs.ollama.com/capabilities/decision>`__ 0.35
and later serves
+ it for the decision models it runs, ``strands-decider serve`` serves
AWS's
+ `Strands Decider <https://github.com/strands-labs/strands-decider>`__,
and
+ `Kev <https://github.com/jaredpalmer/kev>`__ is an open-weight model
built to serve
+ it. `pydantic-ai's System One page
<https://pydantic.dev/docs/ai/models/system-one/>`__
+ lists more.
+
+Both prefixes need a newer ``pydantic-ai-slim`` than the floor the provider's
other extras
+set, so check the release you have installed.
+
Setup
-----
-1. Install the extra:
+The examples on this page use a connection named ``decision_default``. Create
it for the
+backend you run.
+
+**TypeSafe Jev**
+
+1. Install the extra, which adds the TypeSafe SDK:
.. code-block:: bash
pip install 'apache-airflow-providers-common-ai[typesafe]'
- The extra installs the TypeSafe SDK. The ``typesafe:`` model adapter is
part of
- pydantic-ai itself from ``pydantic-ai-slim`` 2.45.0, which is newer than
the floor the
- provider's other extras set, so check that release or later is installed.
-
2. Create a connection (``Admin > Connections``):
- - **Connection Id**: ``jev_default``
+ - **Connection Id**: ``decision_default``
- **Connection Type**: ``Pydantic AI``
- **Password**: your TypeSafe API key, from your `TypeSafe account
<https://typesafe.ai>`__
- **Extra**: ``{"model": "typesafe:jev-1.13.0"}``
-Leave **Host** empty unless you are pointing at a proxy; the provider defaults
to
-TypeSafe's own endpoint.
+ Leave **Host** empty unless you are pointing at a proxy; the provider
defaults to
+ TypeSafe's own endpoint.
+
+**A System One server**
-Pin the version rather than using ``jev-latest``. A threshold you tuned
against one
-release is not guaranteed to mean the same thing after the next one, and
``jev-latest``
-moves under you.
+1. Start the server, or note the URL of a hosted one. For example, to serve
Strands Decider
+ on your own machine:
-When a classifier model is the right choice
--------------------------------------------
+ .. code-block:: bash
+
+ pip install strands-decider
+ strands-decider serve StrandsAgents/strands-decider-2B-hobson-v19
--port 8000
+
+2. Create a connection (``Admin > Connections``):
+
+ - **Connection Id**: ``decision_default``
+ - **Connection Type**: ``Pydantic AI``
+ - **Host**: the server's URL, such as ``http://decider.internal:8000``,
with or without a
+ trailing ``/v1``
+ - **Password**: the server's API key, sent as a bearer token. Leave it
empty for a server
+ that takes none, such as a local Ollama.
+ - **Extra**: ``{"model": "system-one:strands-decider-2B-hobson-v19"}``
+
+ The name after ``system-one:`` is sent to the server as the model to answer
with. A
+ server that runs several, such as Ollama, picks one by it.
+
+3. Describe every option you ask about, and give the question its text.
Servers differ in
+ what they accept, and Strands Decider 0.1.0 refuses, with an HTTP 422 that
fails the task,
+ a question that has an option without a description or that has no question
text at all.
+ Describe each branch in ``branches`` and each member of an ``Enum`` through
a docstring;
+ the members of a bare ``Literal`` have no descriptions, so use a described
``Enum``
+ instead, as the worked example below does. The question text is the
operator's
+ ``system_prompt`` or the agent's ``instructions``, which default to empty.
+
+Pin the model and its version, as in ``typesafe:jev-1.13.0`` rather than
+``typesafe:jev-latest``. A threshold you tuned against one model, or one
release of it, is
+not guaranteed to mean the same thing on another, so measure it again when you
change
+either.
+
+When a decision model is the right choice
+-----------------------------------------
All four of these have to hold.
@@ -70,11 +127,14 @@ All four of these have to hold.
``str`` field is refused, so anything that writes a summary, a query, a
migration, or a
reply to a person is out.
-**You can name the options up front, and there are at most 255.** Downstream
task ids,
-error categories, severity levels, environments, a repository list. If the set
is open, or
-is discovered at run time and could grow past the cap, this is the wrong tool.
The cap is
-per question, so tools attached to an agent get their own 255, with the output
type as one
-option in that question.
+**You can name the options up front, and there are few enough for the model.**
Downstream
+task ids, error categories, severity levels, environments, a repository list.
If the set is
+open, or is discovered at run time and could grow past the cap, this is the
wrong tool. Each
+model has its own cap per question: Jev takes 255 options, and Ollama takes
26. Tools
+attached to an agent count as options in a question of their own, with the
output type as
+one more. pydantic-ai refuses a question over Jev's cap before sending it; a
``system-one:``
+model's cap is the server's, so a question over it comes back as an HTTP error
from the
+server instead.
**The decision is on a path where latency or cost is the constraint.** One
decision per
Dag run rarely justifies changing models. One per row, per file, or per
retrieved document
@@ -88,7 +148,7 @@ itself. If you would not do anything different with a
confidence of 0.55 than wi
that is a sign a general-purpose model is fine here.
And one case where the answer is neither: if a deterministic rule already
sorts the input
-correctly, use the rule. A classifier model is cheap, not free, and a rule you
can read is
+correctly, use the rule. A decision model is cheap, not free, and a rule you
can read is
worth more than a probability you have to calibrate.
Where it fits in this provider
@@ -104,9 +164,10 @@ Where it fits in this provider
* -
:class:`~airflow.providers.common.ai.operators.llm_branch.LLMBranchOperator`
- Yes, with a caveat
- The downstream task ids are already presented to the model as a
constrained set of
- choices, which is exactly the shape a classifier model answers. Setting
- ``model_id`` is the only change, as long as there are two or more
downstream tasks
- -- a one-option pick is refused. Describe each branch in ``branches``
and set a
+ choices, which is exactly the shape a decision model answers. Pointing
+ ``llm_conn_id`` (or ``model_id``) at one is the only change, as long as
there are two
+ or more downstream tasks -- a one-option pick is refused. Describe each
branch in
+ ``branches`` and set a
``decision_policy`` so an unsure pick goes to a person instead of
branching; see
:doc:`operators/llm_branch`.
* - :class:`~airflow.providers.common.ai.operators.llm.LLMOperator` /
@@ -115,7 +176,7 @@ Where it fits in this provider
- Yes
- A ``Literal``, ``Enum``, ``bool`` or bounded number works. Describe the
field, which
becomes the question, and describe each option, which is what tells
them apart. An
- option with no description is read from its name alone.
+ option with no description is read from its name alone, and some
servers refuse it.
* - :doc:`ClassifierRetryPolicy <retry_policies>`
- Yes
- The model names one of the policy's ``categories`` and nothing else;
retry or
@@ -123,12 +184,12 @@ Where it fits in this provider
worker. Set ``min_confidence`` and an unsure answer goes to
``fallback_policy``
(typically an ``LLMRetryPolicy`` on a text model), then
``fallback_rules``, then
the task's own retry behaviour, instead of ending the task on the
model's say-so.
- ``LLMRetryPolicy`` itself asks for free text, which a classifier model
refuses.
+ ``LLMRetryPolicy`` itself asks for free text, which a decision model
refuses.
This is the surface where the model's speed and price matter most: it
runs on
every task failure.
* - Agents with toolsets
- Partly
- - Which tool the text calls for is itself a pick, so a classifier model
can make it.
+ - Which tool the text calls for is itself a pick, so a decision model can
make it.
What it cannot write is a tool's arguments. A tool taking none it calls
itself; one
taking arguments raises ``ToolCallProposed`` after the request, which
is a
``ModelAPIError`` rather than a refusal, so ``FallbackModel`` hands
those requests to
@@ -155,21 +216,35 @@ transcript when that is enabled, and a hook-level call
has it on the result:
.. code-block:: python
+ from enum import Enum
+
+ from pydantic_ai import UseEnumMemberDocstrings
+
from airflow.providers.common.ai.hooks.pydantic_ai import PydanticAIHook
from airflow.sdk import task
- from typing import Literal
+
+
+ class FailureCause(UseEnumMemberDocstrings, str, Enum):
+ transient = "transient"
+ """A fault that clears by itself: a timeout, throttling, a dropped
connection."""
+
+ resource = "resource"
+ """A dependency is down or unreachable and needs fixing before a retry
can work."""
+
+ permanent = "permanent"
+ """A bug or bad input that fails the same way however often it is
retried."""
@task
def triage(log_line: str) -> dict:
- agent = PydanticAIHook(llm_conn_id="jev_default").create_agent(
- output_type=Literal["transient", "resource", "permanent"],
+ agent = PydanticAIHook(llm_conn_id="decision_default").create_agent(
+ output_type=FailureCause,
instructions="Classify why this Airflow task failed.",
)
result = agent.run_sync(log_line)
details = result.response.provider_details or {}
confidence = (details.get("confidence") or {}).get("response")
- return {"category": result.output, "confidence": confidence}
+ return {"category": result.output.value, "confidence": confidence}
Then branch on the returned confidence in a downstream task, so an unsure
answer escalates
instead of acting. Use a higher bar for acting automatically than for flagging
something
@@ -179,21 +254,22 @@ shape for the decision.
Worked example
--------------
-`example_classifier_model.py
<https://github.com/apache/airflow/blob/providers-common-ai/|version|/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_classifier_model.py>`__
-has both halves: a branch whose only classifier-specific line is ``model_id``,
and a
+`example_decision_model.py
<https://github.com/apache/airflow/blob/providers-common-ai/|version|/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_decision_model.py>`__
+has both halves: a branch with nothing specific to decision models but its
connection, and a
classification that escalates when the confidence is low.
What it answers badly
---------------------
-Read `pydantic-ai's model page
<https://pydantic.dev/docs/ai/models/typesafe/>`__ and
-`TypeSafe's own documentation <https://docs.typesafe.ai/>`__ before you
-trust a number from one of these models. Two of its failure modes matter more
than the
-rest in a Dag:
+Read `pydantic-ai's decision model guide
<https://pydantic.dev/docs/ai/models/decision/>`__
+and the documentation of the model you run (`TypeSafe's
<https://docs.typesafe.ai/>`__ for
+Jev) before you trust a number from one of these models. Two failure modes
documented for
+Jev matter more than the rest in a Dag; check them against any other model
before relying on
+it not to share them:
- **The text is treated as data, not as hostile.** An injected instruction, a
misleading
framing, or an argument for its own answer can move the result. A guard
built on a
- classifier model belongs alongside deterministic checks, not instead of them.
+ decision model belongs alongside deterministic checks, not instead of them.
- **Option order is part of what the model sees.** Reordering a ``Literal``'s
members can
change the answer, so a threshold measured against one ordering is measured
against
that ordering only.
diff --git a/providers/common/ai/docs/examples.rst
b/providers/common/ai/docs/examples.rst
index 592000c9938..0672b0d4b67 100644
--- a/providers/common/ai/docs/examples.rst
+++ b/providers/common/ai/docs/examples.rst
@@ -165,14 +165,14 @@ Reliability
- What it shows
* - :doc:`retry_policies`
- Classifying task failures with an LLM into categories you define, then
deriving
- retry, fail, or delay from the category; and the same on a classifier
model with a
+ retry, fail, or delay from the category; and the same on a decision
model with a
confidence bar. Source:
`example_llm_retry_policy.py
<https://github.com/apache/airflow/blob/providers-common-ai/|version|/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_llm_retry_policy.py>`__.
* - :doc:`provider_fallback`
- Failing over to another vendor inside one task attempt, and drilling
the chain
without waiting for an outage. Source:
`example_llm_fallback.py
<https://github.com/apache/airflow/blob/providers-common-ai/|version|/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_llm_fallback.py>`__.
- * - :doc:`classifier_models`
+ * - :doc:`decision_models`
- Routing a failure with a model that answers typed questions instead of
writing
text, and escalating when its confidence is low. Source:
- `example_classifier_model.py
<https://github.com/apache/airflow/blob/providers-common-ai/|version|/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_classifier_model.py>`__.
+ `example_decision_model.py
<https://github.com/apache/airflow/blob/providers-common-ai/|version|/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_decision_model.py>`__.
diff --git a/providers/common/ai/docs/installation.rst
b/providers/common/ai/docs/installation.rst
index 77e8b4e3e1f..a7227e724d6 100644
--- a/providers/common/ai/docs/installation.rst
+++ b/providers/common/ai/docs/installation.rst
@@ -40,8 +40,9 @@ The provider's extras split into a few groups:
* **Model providers** (``openai``, ``anthropic``, ``google``, ``bedrock``,
``typesafe``):
pick the one matching your ``llm_conn_id`` connection
(:doc:`model_providers` maps
vendors to extras, prefixes and connection types). ``typesafe`` differs from
the rest
- in kind: it installs a classifier model that answers typed questions and
cannot write
- text (see :doc:`classifier_models`). The first four mirror the identically
named
+ in kind: it installs a decision model that answers typed questions and
cannot write
+ text (see :doc:`decision_models`). Other decision models, behind the System
One API,
+ need no extra. The first four mirror the identically named
``pydantic-ai-slim`` optional dependency groups, and ``typesafe`` adds the
``typesafe-sdk``
the built-in adapter talks to; pydantic-ai supports more model providers
than these, each under its own extra name, so check the
diff --git a/providers/common/ai/docs/model_providers.rst
b/providers/common/ai/docs/model_providers.rst
index 7e31a70b648..fcfe587903a 100644
--- a/providers/common/ai/docs/model_providers.rst
+++ b/providers/common/ai/docs/model_providers.rst
@@ -92,11 +92,17 @@ type shown, and set the model name with that prefix.
- ``pydanticai``
- ``SNOWFLAKE_ACCOUNT`` and ``SNOWFLAKE_TOKEN`` in the worker
environment; leave
**Password** empty
- * - TypeSafe Jev (classifier, does not write text)
+ * - TypeSafe Jev (decision model, does not write text)
- ``typesafe:``
- ``[typesafe]``
- - ``pydanticai`` (:doc:`classifier_models`)
+ - ``pydanticai`` (:doc:`decision_models`)
- API key in **Password**
+ * - Decision models behind the System One API: Ollama, Strands Decider, Kev
and others
+ (do not write text)
+ - ``system-one:``
+ - None (``pydantic-ai-slim`` 2.53.0+)
+ - ``pydanticai`` with the server URL in **Host** (:doc:`decision_models`)
+ - API key in **Password**, if the server takes one
``[name]`` in the Install column is an extra of this provider, installed as
``pip install "apache-airflow-providers-common-ai[name]"``. The Groq, Mistral
and Snowflake entries
@@ -141,7 +147,7 @@ Pages in this section
AWS Bedrock <connections/pydantic_ai_bedrock>
Google Vertex AI <connections/pydantic_ai_vertex>
Self-hosted models <self_hosted_models>
- Classifier models <classifier_models>
+ Decision models <decision_models>
Provider fallback <provider_fallback>
Connection reference <connections/pydantic_ai>
Using the hook directly <hooks/pydantic_ai>
diff --git a/providers/common/ai/docs/operators/llm_branch.rst
b/providers/common/ai/docs/operators/llm_branch.rst
index 6b07b1949a2..35b15f0f2b9 100644
--- a/providers/common/ai/docs/operators/llm_branch.rst
+++ b/providers/common/ai/docs/operators/llm_branch.rst
@@ -84,8 +84,7 @@ it, so the model can only answer with one of the task IDs.
Descriptions explain the choices; they do not make the model more certain,
and a text model's structured output carries no confidence to read. With a
-classifier model such as TypeSafe's, the descriptions become the criteria of
-its choice question, which is the text it weighs each option by.
+decision model, the descriptions become the criteria of its choice question,
which is the text it weighs each option by.
A pick is relative: the model chooses the best fit among the downstream tasks
offered, not whether any of them fits. If "none of these" or "not enough to
@@ -176,8 +175,7 @@ Reviewing Uncertain Picks
Experimental: this can change or be removed in a minor release of this
provider.
See :ref:`howto/stability`.
-A classifier model such as TypeSafe's returns a confidence with every pick,
-a number from 0 to 1 that summarizes how concentrated its probability
+A decision model returns a confidence with every pick, a number from 0 to 1
that summarizes how concentrated its probability
distribution was: near 1 when one branch stood out, low when two or more
were close. It is not the probability that the pick is right. It is the
model saying how clear-cut the question was, and it is the signal you gate on.
@@ -229,9 +227,9 @@ review it opens is the same one ``require_approval`` opens:
``approval_timeout``
``on_approval_timeout``, ``allow_modifications``, ``approval_notifiers`` and
``approval_assigned_users`` all apply to it.
-TypeSafe's models need the provider's ``typesafe`` extra and a ``pydanticai``
-connection whose Model is ``typesafe:jev-1.13.0`` (or the ``model_id`` on the
-operator, as in the example). :doc:`../classifier_models` covers the setup.
+The example reads a ``pydanticai`` connection, ``decision_default``, whose
Model is a decision
+model: ``typesafe:jev-1.13.0`` for TypeSafe's Jev, or ``system-one:<model>``
for a server
+answering the System One API. :doc:`../decision_models` covers the setup for
each.
The decision record
^^^^^^^^^^^^^^^^^^^
@@ -293,7 +291,7 @@ At execution time, the operator:
3. Passes that type as ``output_type`` to ``pydantic-ai``, constraining the LLM
to valid task IDs only.
4. Reads the model's confidence for the pick from ``provider_details`` (a
- classifier model reports one; a text model does not), pushes the
+ decision model reports one; a text model does not), pushes the
``decision`` XCom, and if the ``decision_policy`` or ``require_approval``
says
so, pauses for human review (or fails, with ``on_uncertain="fail"``).
5. Converts the LLM's structured output to task ID string(s) and calls
diff --git a/providers/common/ai/docs/redirects.txt
b/providers/common/ai/docs/redirects.txt
index 295150c2603..a7c8c571403 100644
--- a/providers/common/ai/docs/redirects.txt
+++ b/providers/common/ai/docs/redirects.txt
@@ -22,3 +22,4 @@ sandbox.rst sandbox/index.rst
hooks/index.rst concepts.rst
end_to_end_pipelines.rst use_cases/index.rst
guardrails.rst capabilities.rst
+classifier_models.rst decision_models.rst
diff --git a/providers/common/ai/docs/retry_policies.rst
b/providers/common/ai/docs/retry_policies.rst
index ace62c46990..79625373783 100644
--- a/providers/common/ai/docs/retry_policies.rst
+++ b/providers/common/ai/docs/retry_policies.rst
@@ -42,8 +42,7 @@ fully model-driven, with the SDK's ``ExceptionRetryPolicy``
as the bottom rung:
(``ClassifierRetryPolicy``)
- The model names one of your ``categories``; the table says whether that
category is retried, after how long, and how sure the model has to be.
- A classifier model such as TypeSafe's Jev answers in a few hundred
- milliseconds and reports its confidence; a text model can sit here too,
+ A decision model answers with its confidence; a text model can sit here
too,
without a bar.
- Category descriptions and the confidence bar. No reasoning, no prose.
* - **LLM**
@@ -174,7 +173,7 @@ failures belong there (the ``description`` the model
reads), whether it is
retried, after what ``delay``, and how sure the model has to be
(``min_confidence``, covered below). Everything the model is told about a
category, and everything the policy does with it, sits in that one entry, so
-the two cannot drift apart. This is also the policy a classifier model needs:
+the two cannot drift apart. This is also the policy a decision model needs:
such a model refuses the free-text fields of ``ErrorClassification``, so an
``LLMRetryPolicy`` pointed at one fails every classification and falls back,
with a log line saying to use ``ClassifierRetryPolicy``.
@@ -287,7 +286,7 @@ constructed, at Dag parse time, rather than on the first
task failure.
Confidence
----------
-A classifier model reports how sure it is of its answer. ``min_confidence`` is
+A decision model reports how sure it is of its answer. ``min_confidence`` is
the bar that answer needs for the policy to act on it; under the bar the policy
discards the answer and takes the same path it takes when the model call fails:
``fallback_rules`` if one matches, otherwise the task's own retry behaviour. It
@@ -296,7 +295,7 @@ does not substitute a delay of its own.
Each category can carry its own bar. The stakes differ: a wrong ``transient``
costs one more attempt, while a wrong ``permanent`` costs the task every retry
it
had left, so the category that ends the task deserves the higher bar.
-``jev_default`` is the classifier-model connection from
:doc:`classifier_models`.
+``decision_default`` is the decision-model connection from
:doc:`decision_models`.
.. exampleinclude::
/../../ai/src/airflow/providers/common/ai/example_dags/example_llm_retry_policy.py
:language: python
@@ -311,7 +310,7 @@ logs the confidence.
confidence, and so does a response whose metadata was dropped along the way.
With no bar configured that changes nothing. With a bar configured, every such
answer is discarded and the fallback path decides, so swapping the connection
-from a classifier model to a text model does not silently switch off a control
+from a decision model to a text model does not silently switch off a control
you set on purpose. To run a text model, remove the bar.
The confidence is a statistic on the shape of the probability distribution the
@@ -323,7 +322,7 @@ answers start. A bar reduces wrong actions and does not
eliminate them: a wrong
pick can arrive with high confidence. Pin the model version
(``typesafe:jev-1.13.0``, not ``jev-latest``): a bar tuned against one release
is not guaranteed to mean the same thing after the next. See
-:doc:`classifier_models` for what these models answer well and badly.
+:doc:`decision_models` for what these models answer well and badly.
Escalating to an LLM
--------------------
@@ -507,7 +506,7 @@ When writing custom instructions:
come back: a model that insists on one is re-prompted once by pydantic-ai and
then gives up, which lands the task on ``fallback_rules`` or on its own retry
behaviour, having billed two calls.
-- A classifier model sends ``instructions`` as the question it scores the
+- A decision model sends ``instructions`` as the question it scores the
exception text against, not as rules it follows step by step, so a long
rubric
buys less there than a better description on each category does.
diff --git a/providers/common/ai/docs/self_hosted_models.rst
b/providers/common/ai/docs/self_hosted_models.rst
index ef1244f78f2..f454d0a2a2b 100644
--- a/providers/common/ai/docs/self_hosted_models.rst
+++ b/providers/common/ai/docs/self_hosted_models.rst
@@ -374,6 +374,9 @@ Where to go next
-------------------
- :ref:`howto/connection:pydanticai` -- the full connection field reference.
+- :doc:`decision_models` -- a self-hosted decision model, such as Strands
Decider or one
+ Ollama runs, takes ``system-one:<model>`` with the server URL in ``host``
rather than
+ the ``openai:`` or ``ollama:`` prefixes on this page.
- :doc:`retry_policies` -- the "Local LLM support" section covers pointing
``LLMRetryPolicy`` at a self-hosted endpoint.
- :doc:`examples` -- more runnable Dags against the ``pydanticai``
diff --git a/providers/common/ai/docs/stability.rst
b/providers/common/ai/docs/stability.rst
index 100ba376555..2101c651253 100644
--- a/providers/common/ai/docs/stability.rst
+++ b/providers/common/ai/docs/stability.rst
@@ -111,7 +111,7 @@ Everything this provider ships that is not in the table
above is experimental.
* - Feature
- Why it is experimental
* -
:class:`~airflow.providers.common.ai.policies.retry.ClassifierRetryPolicy`
- (:doc:`classifier_models`)
+ (:doc:`decision_models`)
- The confidence threshold, the fallback order and the behaviour when the
classifier
is unavailable are still settling.
* - :class:`~airflow.providers.common.ai.policies.decision.DecisionPolicy`,
and
diff --git a/providers/common/ai/docs/use_cases/index.rst
b/providers/common/ai/docs/use_cases/index.rst
index 027d79c33e4..7793e4e4c5d 100644
--- a/providers/common/ai/docs/use_cases/index.rst
+++ b/providers/common/ai/docs/use_cases/index.rst
@@ -30,7 +30,8 @@ the shape most of the others build on: structured output plus
dynamic task mappi
Every Dag needs the provider installed with the extra for your model vendor
and a
``pydanticai`` connection named ``pydanticai_default``; :doc:`../quickstart`
covers both.
-Each page's "Run it" lists only what that Dag adds.
+Each page's "Run it" lists only what that Dag adds.
:doc:`route_pipeline_failures` is the
+exception: it reads a decision-model connection, ``decision_default``, instead.
.. list-table::
:header-rows: 1
diff --git a/providers/common/ai/docs/use_cases/route_pipeline_failures.rst
b/providers/common/ai/docs/use_cases/route_pipeline_failures.rst
index 046961a8cbf..10ca73f304a 100644
--- a/providers/common/ai/docs/use_cases/route_pipeline_failures.rst
+++ b/providers/common/ai/docs/use_cases/route_pipeline_failures.rst
@@ -28,23 +28,18 @@ What this demonstrates
----------------------
* :doc:`../operators/llm_branch` -- ``LLMBranchOperator`` picks one downstream
task id.
-* :doc:`../classifier_models` -- a classifier model, so the gate has a
confidence to read.
+* :doc:`../decision_models` -- a decision model, so the gate has a confidence
to read.
* :doc:`../approval_gates` -- ``on_uncertain="review"`` routes low-confidence
picks to a
human, with ``approval_timeout`` bounding the wait.
Run it
------
-1. Install the provider with the classifier extra:
+1. Create the ``decision_default`` connection for the decision model you run,
TypeSafe Jev or
+ a server answering the System One API (:doc:`../decision_models` covers
both). Both Dags
+ below read it.
- .. code-block:: bash
-
- pip install "apache-airflow-providers-common-ai[typesafe]"
-
-2. Set ``{"model": "typesafe:jev-1.13.0"}`` in the ``pydanticai_default``
extra. The second
- Dag below reads the same kind of connection as ``jev_default``.
-
-3. Trigger the Dag:
+2. Trigger the Dag:
.. code-block:: bash
@@ -67,13 +62,13 @@ Classify-then-act variant
-------------------------
When the action depends on the score itself, classify in one task and act in
the next
-(:doc:`../classifier_models` explains reading the score). It uses the
``jev_default``
-connection; run it as ``airflow dags test
example_classifier_model_confidence``:
+(:doc:`../decision_models` explains reading the score). It uses the same
``decision_default``
+connection; run it as ``airflow dags test example_decision_model_confidence``:
-.. exampleinclude::
/../../ai/src/airflow/providers/common/ai/example_dags/example_classifier_model.py
+.. exampleinclude::
/../../ai/src/airflow/providers/common/ai/example_dags/example_decision_model.py
:language: python
- :start-after: [START howto_classifier_model_confidence]
- :end-before: [END howto_classifier_model_confidence]
+ :start-after: [START howto_decision_model_confidence]
+ :end-before: [END howto_decision_model_confidence]
Adapting it
-----------
diff --git
a/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_classifier_model.py
b/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_decision_model.py
similarity index 60%
rename from
providers/common/ai/src/airflow/providers/common/ai/example_dags/example_classifier_model.py
rename to
providers/common/ai/src/airflow/providers/common/ai/example_dags/example_decision_model.py
index b56124f50fc..179d7bf3066 100644
---
a/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_classifier_model.py
+++
b/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_decision_model.py
@@ -15,35 +15,41 @@
# specific language governing permissions and limitations
# under the License.
"""
-Example DAG using a classifier model to route an incident.
+Example DAG using a decision model to route an incident.
-A classifier model answers typed questions and cannot write text, so it suits
a branch
+A decision model answers typed questions and cannot write text, so it suits a
branch
(pick one of these task ids) and a classification (pick one of these labels),
and nothing
-that produces prose. See the "Classifier models" guide in this provider's
documentation.
+that produces prose. See the "Decision models" guide in this provider's
documentation.
-Prerequisites:
- - ``pip install 'apache-airflow-providers-common-ai[typesafe]'``
- - Connection ``jev_default`` with ``conn_type='pydanticai'``,
``password=<API key>``,
- ``extra='{"model": "typesafe:jev-1.13.0"}'``
+Prerequisites: ``pydantic-ai-slim`` 2.46.0 or later (2.53.0 for
``system-one:``), and a
+connection ``decision_default`` with ``conn_type='pydanticai'`` whose model is
a decision model,
+either:
+
+- TypeSafe Jev: ``password=<API key>``, ``extra='{"model":
"typesafe:jev-1.13.0"}'``, and
+ ``pip install 'apache-airflow-providers-common-ai[typesafe]'``, or
+- a server answering the System One API: ``host=<server URL>``,
``password=<API key, if any>``,
+ ``extra='{"model": "system-one:<model name>"}'``
"""
from __future__ import annotations
-from typing import Literal
+from enum import Enum
+
+from pydantic_ai import UseEnumMemberDocstrings
from airflow.providers.common.ai.hooks.pydantic_ai import PydanticAIHook
from airflow.providers.common.ai.operators.llm_branch import LLMBranchOperator
from airflow.providers.common.compat.sdk import dag, task
# Clear-cut: the error names the missing permission, so one remediation fits
and the others
-# do not. Branch options a classifier model can separate are the ones to give
it.
+# do not. Branch options a decision model can separate are the ones to give it.
PERMISSION_FAILURE = (
"Task export_report failed: botocore.exceptions.ClientError: An error
occurred "
"(AccessDenied) when calling the PutObject operation: the task role "
"airflow-worker lacks s3:PutObject on arn:aws:s3:::reports-prod/daily/."
)
-# Answerable but not decisively: a classifier model lands on ``resource`` here
with roughly
+# Answerable but not decisively: a decision model lands on ``resource`` here
with roughly
# two thirds of the probability, which is enough to be worth recording and not
enough to act
# on unattended. That is the case the confidence gate below exists for.
UNDERDETERMINED_INCIDENT = (
@@ -59,16 +65,22 @@ ACT_ABOVE = 0.8
REVIEW_ABOVE = 0.5
-# [START howto_classifier_model_branch]
-@dag(tags=["example", "classifier"])
-def example_classifier_model_branch():
- """Route the incident. The model id is the only thing that makes this a
classifier."""
+# [START howto_decision_model_branch]
+@dag(tags=["example", "decision"])
+def example_decision_model_branch():
+ """Route the incident. The connection's model is the only thing that makes
this a decision model."""
route = LLMBranchOperator(
task_id="route_failure",
prompt=PERMISSION_FAILURE,
- llm_conn_id="jev_default",
- model_id="typesafe:jev-1.13.0",
+ llm_conn_id="decision_default",
system_prompt="Pick the remediation that addresses the cause, not the
symptom.",
+ # Each description is what the model weighs that branch by. Some
System One servers
+ # refuse an option without one, so describe every branch.
+ branches={
+ "grant_bucket_write": "The task's role lacks a permission it needs
on the bucket.",
+ "restore_deleted_bucket": "The destination bucket or prefix no
longer exists.",
+ "wait_and_retry": "A transient fault: throttling, a timeout, a
brief outage.",
+ },
)
@task
@@ -86,14 +98,29 @@ def example_classifier_model_branch():
route >> [grant_bucket_write(), restore_deleted_bucket(), wait_and_retry()]
-# [END howto_classifier_model_branch]
+# [END howto_decision_model_branch]
+
+example_decision_model_branch()
+
+
+# [START howto_decision_model_confidence]
+class FailureCause(UseEnumMemberDocstrings, str, Enum):
+ """Why an Airflow task failed."""
+
+ # Each docstring is what the model weighs that option by. Some System One
servers
+ # refuse an option without one, so describe every option.
+ transient = "transient"
+ """A fault that clears by itself: a timeout, throttling, a dropped
connection."""
+
+ resource = "resource"
+ """A dependency is down or unreachable and needs fixing before a retry can
work."""
-example_classifier_model_branch()
+ permanent = "permanent"
+ """A bug or bad input that fails the same way however often it is
retried."""
-# [START howto_classifier_model_confidence]
-@dag(tags=["example", "classifier"])
-def example_classifier_model_confidence():
+@dag(tags=["example", "decision"])
+def example_decision_model_confidence():
"""Classify a failure and escalate when the model says it does not know.
``LLMBranchOperator`` can gate its pick on one confidence bar through
``decision_policy``.
@@ -103,8 +130,8 @@ def example_classifier_model_confidence():
@task
def classify(log_line: str) -> dict:
- agent = PydanticAIHook(llm_conn_id="jev_default").create_agent(
- output_type=Literal["transient", "resource", "permanent"],
+ agent = PydanticAIHook(llm_conn_id="decision_default").create_agent(
+ output_type=FailureCause,
instructions="Classify why this Airflow task failed.",
)
result = agent.run_sync(log_line)
@@ -115,7 +142,7 @@ def example_classifier_model_confidence():
# same confidence in their ``decision`` XCom.
details = result.response.provider_details or {}
return {
- "category": result.output,
+ "category": result.output.value,
"confidence": (details.get("confidence") or {}).get("response"),
}
@@ -131,6 +158,6 @@ def example_classifier_model_confidence():
act(classify(UNDERDETERMINED_INCIDENT))
-# [END howto_classifier_model_confidence]
+# [END howto_decision_model_confidence]
-example_classifier_model_confidence()
+example_decision_model_confidence()
diff --git
a/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_llm_branch.py
b/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_llm_branch.py
index 0a17f9c43a1..d210395acc9 100644
---
a/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_llm_branch.py
+++
b/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_llm_branch.py
@@ -101,7 +101,7 @@ example_llm_branch_descriptions()
# [START howto_operator_llm_branch_decision_policy]
@dag(tags=["example"])
def example_llm_branch_decision_policy():
- # A classifier model reports how sure it is of each pick; a text model
does not, and
+ # A decision model reports how sure it is of each pick; a text model does
not, and
# with a min_confidence set every pick would count as uncertain and go to
review.
route = LLMBranchOperator(
task_id="triage_failure",
@@ -109,8 +109,7 @@ def example_llm_branch_decision_policy():
"Task load_orders failed: psycopg2.OperationalError: could not
connect to server: "
"Connection timed out. Is the server running on host db.internal
(10.0.4.12)?"
),
- llm_conn_id="pydanticai_default",
- model_id="typesafe:jev-1.13.0",
+ llm_conn_id="decision_default",
system_prompt="Pick the remediation that addresses the cause of the
failure.",
branches={
"rerun": "The failure looks transient: a timeout, a dropped
connection, a rate limit.",
diff --git
a/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_llm_retry_policy.py
b/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_llm_retry_policy.py
index 869fee2e549..a2dd6a52444 100644
---
a/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_llm_retry_policy.py
+++
b/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_llm_retry_policy.py
@@ -19,16 +19,16 @@ Example DAG demonstrating LLM-powered retry policies.
``llm_policy`` is the plain form: a text model classifies the failure and
decides whether to
retry and how long to wait, guided by its instructions. ``patient_policy`` and
the
-classifier-model policies are ``ClassifierRetryPolicy``: the model only names
the kind of failure
+decision-model policies are ``ClassifierRetryPolicy``: the model only names
the kind of failure
and the policy's table decides the action, the delay and how sure the model
has to be.
Prerequisites:
- Connection ``pydanticai_default`` with ``conn_type='pydanticai'``,
``password=<API key>``, ``extra='{"model":
"anthropic:claude-haiku-4-5-20251001"}'``
- ``pip install apache-airflow-providers-common-ai[anthropic]``
- - For the classifier-model Dag: connection ``jev_default`` with
- ``extra='{"model": "typesafe:jev-1.13.0"}'`` and
- ``pip install 'apache-airflow-providers-common-ai[typesafe]'``
+ - For the decision-model Dag: a connection ``decision_default`` whose model
is a decision
+ model, as set up in the "Decision models" guide. Its confidence bars were
calibrated on
+ ``typesafe:jev-1.13.0``; measure them again for another model.
"""
from __future__ import annotations
@@ -96,15 +96,15 @@ try:
example_llm_retry_policy()
# [START howto_retry_policy_classifier]
- # A classifier model answers the same question in a few hundred
milliseconds and reports
- # how sure it is. The categories are this pipeline's own; their
descriptions are what the
+ # A decision model answers the same question with how sure it is, and
jev-1.13.0 does it
+ # in a few hundred milliseconds. The categories are this pipeline's own;
their descriptions are what the
# model reads. ``permanent`` ends the task and costs it every retry it had
left, so it
# demands more certainty than the rest. Under a bar the answer is
discarded and the
# fallback rules, then the task's own retry settings, decide instead. The
bars come from
# a calibration run on jev-1.13.0: correct picks landed at 0.89 and above,
wrong ones
# at 0.47 to 0.69, with one wrong ``permanent`` at 0.90 that no sensible
bar catches.
snowflake_policy = ClassifierRetryPolicy(
- llm_conn_id="jev_default",
+ llm_conn_id="decision_default",
min_confidence=0.8,
categories={
"queued": ErrorCategory(
@@ -134,11 +134,11 @@ try:
# [END howto_retry_policy_classifier]
# [START howto_retry_policy_escalation]
- # The three layers chained. The classifier answers the clear cases in a
few hundred
- # milliseconds. When it is unsure or unreachable, a text model reasons
about the failure
- # and decides retry and delay itself. When that model is unreachable too,
the rules decide.
+ # The three layers chained. The decision model answers the clear cases
cheaply. When it is
+ # unsure or unreachable, a text model reasons about the failure and
decides retry and delay
+ # itself. When that model is unreachable too, the rules decide.
escalating_policy = ClassifierRetryPolicy(
- llm_conn_id="jev_default",
+ llm_conn_id="decision_default",
min_confidence=0.8,
categories=snowflake_policy.categories,
fallback_policy=LLMRetryPolicy(llm_conn_id="pydanticai_default",
timeout=30.0),
diff --git
a/providers/common/ai/src/airflow/providers/common/ai/operators/llm.py
b/providers/common/ai/src/airflow/providers/common/ai/operators/llm.py
index f2ce27a315a..3c11b3bd1a8 100644
--- a/providers/common/ai/src/airflow/providers/common/ai/operators/llm.py
+++ b/providers/common/ai/src/airflow/providers/common/ai/operators/llm.py
@@ -158,10 +158,10 @@ class LLMOperator(CancellableAgentRunMixin, BaseOperator,
LLMApprovalMixin):
saying how confident the model has to be for the operator to return
its answer by
itself (``min_confidence``) and what happens otherwise
(``on_uncertain``: ``"review"``
or ``"fail"``). Confidence comes from models that report one per
output field, such as
- a classifier model (TypeSafe's), in ``provider_details``. A structured
output is judged
- by its least confident field among the fields that reported one; a
field whose type
- reports none (a bounded float, where the probability is the answer) is
not gated, and
- the record's ``confidence`` map shows which fields were compared. When
no field reports
+ a decision model, in ``provider_details``. A structured output is
judged by its least
+ confident field among the fields that reported one; a field whose type
reports none (a
+ bounded float, where the probability is the answer) is not gated, and
the record's
+ ``confidence`` map shows which fields were compared. When no field
reports
any confidence, as with a text model, the output counts as uncertain,
so swapping the
connection does not silently switch off a control the author set.
Independent of
``require_approval``, which always asks. ``on_uncertain="review"``
needs Airflow 3.1+,
diff --git
a/providers/common/ai/src/airflow/providers/common/ai/operators/llm_branch.py
b/providers/common/ai/src/airflow/providers/common/ai/operators/llm_branch.py
index 02f603ba3af..2942b1601d4 100644
---
a/providers/common/ai/src/airflow/providers/common/ai/operators/llm_branch.py
+++
b/providers/common/ai/src/airflow/providers/common/ai/operators/llm_branch.py
@@ -98,19 +98,19 @@ class LLMBranchOperator(LLMOperator, BranchMixIn):
output schema next to the option, so the model reads "here is an
option, here is
what it means" rather than guessing from the task ID;
``min_confidence`` on an
option, which is experimental, is a bar for that branch alone. A
downstream task
- without an entry is
- presented by its ID alone and takes the policy's bar. A key that is
not a
- downstream task ID fails the task before the model is called.
Descriptions
- support Jinja templating.
+ without an entry is presented by its ID alone and takes the policy's
bar; some
+ decision-model servers refuse an option without a description, so
describe every
+ branch when the connection is a decision model. A key that is not a
downstream task
+ ID fails the task before the model is called. Descriptions support
Jinja templating.
:param allow_multiple_branches: When ``False`` (default) the LLM returns a
single task ID. When ``True`` the LLM may return one or more task IDs.
:param decision_policy: Experimental. A
:class:`~airflow.providers.common.ai.policies.decision.DecisionPolicy`:
the confidence a pick needs for the operator to branch on it without a
person
(``min_confidence``) and what happens under it (``on_uncertain``:
``"review"`` or
- ``"fail"``). Confidence comes from models that report one, such as a
classifier
- model (TypeSafe's); a text model reports none, which counts as
uncertain, so
- swapping the connection does not silently switch off a control the
author set.
+ ``"fail"``). Confidence comes from models that report one, such as a
decision model;
+ a text model reports none, which counts as uncertain, so swapping the
connection does
+ not silently switch off a control the author set.
A branch's own ``min_confidence`` overrides the policy's for that
pick; with
``allow_multiple_branches`` the strictest bar among the picked
branches applies.
``require_approval=True`` still sends every pick to a person
regardless.
diff --git
a/providers/common/ai/src/airflow/providers/common/ai/policies/decision.py
b/providers/common/ai/src/airflow/providers/common/ai/policies/decision.py
index 4e60b2bc2ad..68494c82238 100644
--- a/providers/common/ai/src/airflow/providers/common/ai/policies/decision.py
+++ b/providers/common/ai/src/airflow/providers/common/ai/policies/decision.py
@@ -56,8 +56,8 @@ class DecisionPolicy:
:param min_confidence: The confidence, from 0 to 1, the answer needs for
the operator to act
without a person. ``None`` (default) is no gate: the operator behaves
as it always has.
- Confidence comes from models that report one, such as a classifier
model (TypeSafe's),
- in ``provider_details``. A model that reports none counts as under the
bar, so swapping
+ Confidence comes from models that report one, such as a decision
model, in
+ ``provider_details``. A model that reports none counts as under the
bar, so swapping
the connection to a text model does not silently switch off a control
the author set.
:param on_uncertain: What happens under the bar. ``"review"`` (default)
sends the answer to
human review through the same approval flow as ``require_approval``,
which needs Airflow
diff --git
a/providers/common/ai/src/airflow/providers/common/ai/policies/retry.py
b/providers/common/ai/src/airflow/providers/common/ai/policies/retry.py
index ad21d52b569..6440dcbaec6 100644
--- a/providers/common/ai/src/airflow/providers/common/ai/policies/retry.py
+++ b/providers/common/ai/src/airflow/providers/common/ai/policies/retry.py
@@ -23,8 +23,7 @@ Model-backed retry policies, one per layer of a ladder from
hardcoded to reasoni
* **Classifier.** :class:`ClassifierRetryPolicy`: the model names one of the
author's
``categories`` and the :class:`ErrorCategory` table decides whether that
category is
retried, after how long, and how sure the model has to be. Tuned through
descriptions and
- a confidence bar, not through reasoning. A classifier model such as
TypeSafe's Jev runs
- here; a text model can too.
+ a confidence bar, not through reasoning. A decision model runs here; a text
model can too.
* **LLM.** :class:`LLMRetryPolicy`: a text model classifies the failure,
decides whether to
retry and how long to wait from ``instructions``, and explains itself.
@@ -347,9 +346,8 @@ class LLMRetryPolicy(_ModelRetryPolicy):
Ollama, etc.) for error classification with structured output. The model
returns an :class:`ErrorClassification`: which category the error is,
whether
to retry, how long to wait, and why, all steered by ``instructions``. This
is
- the reasoning layer; for a cheap typed decision from a classifier model
such as
- TypeSafe's Jev, use :class:`ClassifierRetryPolicy`, which can name this
policy as
- its ``fallback_policy``.
+ the reasoning layer; for a cheap typed decision from a decision model, use
+ :class:`ClassifierRetryPolicy`, which can name this policy as its
``fallback_policy``.
When the LLM call itself fails, the policy falls back to ``fallback_rules``
(if provided) or returns DEFAULT to use the task's standard retry logic.
@@ -398,11 +396,11 @@ class LLMRetryPolicy(_ModelRetryPolicy):
result = self._run(agent, exception, try_number, max_tries)
except Exception as exc:
if "not supported by this model" in str(exc):
- # A classifier model refuses ErrorClassification's free-text
fields client-side. A text
+ # A decision model refuses ErrorClassification's free-text
fields client-side. A text
# model whose profile lacks structured output raises the same
words, so this is a hint.
log.error(
- "This model cannot answer ErrorClassification. If it is a
classifier model such as "
- "TypeSafe's Jev, use ClassifierRetryPolicy, which asks it
a typed question."
+ "This model cannot answer ErrorClassification. If it is a
decision model, "
+ "use ClassifierRetryPolicy, which asks it a typed
question."
)
raise
classification = result.output
@@ -441,8 +439,7 @@ class ClassifierRetryPolicy(_ModelRetryPolicy):
The model's only job is to pick one of ``categories``; it reads each one's
description
from the output schema. Whether that category is retried, after how long,
and how sure
the model has to be all come from the :class:`ErrorCategory` in the worker
process.
- That is the shape a classifier model such as TypeSafe's Jev answers, in a
few hundred
- milliseconds and with a confidence; a text model answers it too.
+ That is the shape a decision model answers, with a confidence; a text
model answers it too.
When the model call fails, or the answer is under its confidence bar, the
policy
consults ``fallback_policy`` if set, then ``fallback_rules``, then returns
DEFAULT to use
@@ -467,7 +464,7 @@ class ClassifierRetryPolicy(_ModelRetryPolicy):
outside them is rejected before the policy acts on it.
:param min_confidence: The confidence, from 0 to 1, the model's answer
needs for the
policy to act on it. ``None`` (default) is no bar: the answer is acted
on whatever
- the confidence. Confidence comes from models that report one, such as
a classifier
+ the confidence. Confidence comes from models that report one, such as
a decision
model, in ``provider_details``. Under the bar, or when a bar is set
and the model
reported no confidence, the answer is discarded and
``fallback_policy``, then
``fallback_rules``, then the task's own retry behaviour apply, so
swapping the
diff --git
a/providers/common/ai/src/airflow/providers/common/ai/utils/decision.py
b/providers/common/ai/src/airflow/providers/common/ai/utils/decision.py
index 6d7a94d49a3..32193d1c2fd 100644
--- a/providers/common/ai/src/airflow/providers/common/ai/utils/decision.py
+++ b/providers/common/ai/src/airflow/providers/common/ai/utils/decision.py
@@ -17,9 +17,9 @@
"""
Confidence gating and the decision record shared by the LLM operators.
-A classifier model such as TypeSafe's reports a confidence per output field in
-``ModelResponse.provider_details`` (``{"confidence": {field: float},
"probabilities":
-{field: {option: float}}}``). A text model reports nothing there. These
helpers read
+A decision model reports a confidence per output field in
``ModelResponse.provider_details``
+(``{"confidence": {field: float}, "probabilities": {field: {option:
float}}}``).
+A text model reports nothing there. These helpers read
that, decide whether an answer should go to a person before the operator acts
on it,
and shape the ``decision`` XCom record that says what was proposed, what was
done, and
why.
@@ -71,7 +71,7 @@ __all__ = [
DECISION_XCOM_KEY = "decision"
BARE_OUTPUT_FIELD = "response"
-"""The field name a classifier model reports a bare (non-object) output type's
confidence under."""
+"""The field name a decision model reports a bare (non-object) output type's
confidence under."""
ReviewReason = Literal["require_approval", "below_threshold",
"missing_confidence"]
DecidedBy = Literal["model", "human", "timeout_default", "policy", "timeout"]
@@ -236,7 +236,7 @@ def check_uncertain_action(
raise LowConfidenceError(
f"min_confidence={threshold:.2f} is set for {what} but the model
reported no confidence for "
"its answer, and on_uncertain='fail'. A text model reports none; a
bounded float field "
- "reports none because the probability is the answer. Use a
classifier model, remove the "
+ "reports none because the probability is the answer. Use a
decision model, remove the "
"bar, or set on_uncertain='review'."
)
raise LowConfidenceError(
@@ -311,7 +311,7 @@ def described_choices(name: str, options: Mapping[str, str
| None]) -> type[Any]
With a description on any option the schema renders as ``anyOf`` of
``{const, description}``
instead of a bare ``enum`` list. That is the one JSON Schema shape that
carries a description
- per value, and it is what both a text model's tool schema and
pydantic-ai's TypeSafe adapter
+ per value, and it is what both a text model's tool schema and
pydantic-ai's decision models
read an option's meaning from. On pydantic-ai 2.46+ the type is its
``Choices``; before that,
an ``Enum`` whose schema hook emits the same shape. Either way the model
has to answer with one
of the keys, in the order given, and :func:`picked_key` returns that key
whichever type answered.
diff --git
a/providers/common/ai/tests/unit/common/ai/operators/test_llm_branch.py
b/providers/common/ai/tests/unit/common/ai/operators/test_llm_branch.py
index 0ea1dd6cb30..5ecb7d26abe 100644
--- a/providers/common/ai/tests/unit/common/ai/operators/test_llm_branch.py
+++ b/providers/common/ai/tests/unit/common/ai/operators/test_llm_branch.py
@@ -211,7 +211,7 @@ class TestLLMBranchOperator:
"""The option order the model sees is sorted, not whatever order
downstream_task_ids iterates in.
``downstream_task_ids`` is a set, so its iteration order depends on
string hashing and
- differs between worker processes. Option order is part of the question
for a classifier
+ differs between worker processes. Option order is part of the question
for a decision
model, so it has to be the same on every worker. A reverse-sorted list
stands in for an
unlucky set order; without ``sorted()`` the enum comes out reversed
and this fails.
"""
diff --git a/providers/common/ai/tests/unit/common/ai/policies/test_retry.py
b/providers/common/ai/tests/unit/common/ai/policies/test_retry.py
index 5168b6b1c4f..30ca9d2a679 100644
--- a/providers/common/ai/tests/unit/common/ai/policies/test_retry.py
+++ b/providers/common/ai/tests/unit/common/ai/policies/test_retry.py
@@ -1220,7 +1220,7 @@ class TestLLMRetryPolicy:
@patch(HOOK, autospec=True)
def test_classifier_refusal_logs_the_categories_hint(self, mock_hook_cls,
caplog):
- """A classifier model refuses ErrorClassification's text fields; the
log says what to do."""
+ """A decision model refuses ErrorClassification's text fields; the log
says what to do."""
agent = _install(mock_hook_cls, MagicMock(spec=Agent))
agent.run_sync.side_effect = RuntimeError("Output field 'reasoning' is
not supported by this model")
policy = LLMRetryPolicy(llm_conn_id="test",
model_id="typesafe:jev-1.13.0")