This is an automated email from the ASF dual-hosted git repository.
kaxil pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/airflow.git
The following commit(s) were added to refs/heads/main by this push:
new ac147d07ffe Remove `code_mode` from `AgentOperator` in favor of the
`CodeMode` capability (#74312)
ac147d07ffe is described below
commit ac147d07ffe3f1d390a2fbc062be80000d39e480
Author: Kaxil Naik <[email protected]>
AuthorDate: Tue Oct 6 08:28:16 2026 +0100
Remove `code_mode` from `AgentOperator` in favor of the `CodeMode`
capability (#74312)
code_mode=True only appended a bare CodeMode() to the agent's capabilities,
which capabilities= now does directly. The flag hid CodeMode's own
arguments,
most importantly max_tool_calls, whose default of 100 nested calls per
run_code snippet is too low for agents that fan out over a list.
---
providers/common/ai/docs/capabilities.rst | 6 +-
providers/common/ai/docs/changelog.rst | 10 +++
providers/common/ai/docs/code_mode.rst | 32 ++++++--
providers/common/ai/docs/features.rst | 4 +-
providers/common/ai/docs/index.rst | 2 +-
providers/common/ai/docs/operators/agent.rst | 7 +-
providers/common/ai/docs/sandbox/index.rst | 2 +-
providers/common/ai/docs/stability.rst | 4 +-
providers/common/ai/docs/tool_approval.rst | 8 +-
providers/common/ai/docs/troubleshooting.rst | 10 +--
providers/common/ai/pyproject.toml | 7 +-
.../common/ai/example_dags/example_agent.py | 9 ++-
.../airflow/providers/common/ai/operators/agent.py | 67 +++-------------
.../tests/unit/common/ai/operators/test_agent.py | 89 +---------------------
.../ai/operators/test_agent_tool_approval.py | 1 -
uv.lock | 2 +-
16 files changed, 78 insertions(+), 182 deletions(-)
diff --git a/providers/common/ai/docs/capabilities.rst
b/providers/common/ai/docs/capabilities.rst
index eeca7d716fd..0284396dff6 100644
--- a/providers/common/ai/docs/capabilities.rst
+++ b/providers/common/ai/docs/capabilities.rst
@@ -112,9 +112,9 @@ again. Whether a capability's work is replayed depends on
where it runs:
and capabilities from an agent spec file
- Tools run again. Pass tools you need replayed in ``toolsets=`` instead.
* - pydantic-ai-harness ``CodeMode``
- - Not allowed: the operator raises ``ValueError``, as it does for
``code_mode=True``. This
- includes a ``CodeMode`` inside a ``CombinedCapability`` or a wrapper
such as
- ``PrefixTools``, but not one a capability function builds when the run
starts.
+ - Not allowed: the operator raises ``ValueError``. This includes a
``CodeMode`` inside a
+ ``CombinedCapability`` or a wrapper such as ``PrefixTools``, but not
one a capability
+ function builds when the run starts.
See :doc:`durable_execution` for how the cache works.
diff --git a/providers/common/ai/docs/changelog.rst
b/providers/common/ai/docs/changelog.rst
index 96b4ea26f35..330b23aa439 100644
--- a/providers/common/ai/docs/changelog.rst
+++ b/providers/common/ai/docs/changelog.rst
@@ -51,6 +51,16 @@ Changelog
result or ``result["sample_columns"]`` on a summary. A summary carries no
``columns`` key, so
``result["columns"]`` raises ``KeyError`` on any table wide enough to be
summarized.
+.. note::
+ ``AgentOperator`` no longer accepts ``code_mode``. Pass the
pydantic-ai-harness capability
+ instead: replace ``code_mode=True`` with ``capabilities=[CodeMode()]``, using
+ ``from pydantic_ai_harness import CodeMode``. The capability takes its own
arguments, such
+ as ``max_tool_calls``. If the task also passes
``agent_params={"capabilities": [...]}``,
+ move those into ``capabilities=`` too, since the two cannot be combined. The
import runs
+ when the Dag file is parsed, so the ``code-mode`` extra is now needed by the
Dag processor
+ as well as the workers. The extra's floor is now
``pydantic-ai-harness>=0.24.0``, the first
+ release with ``max_tool_calls``. See :doc:`code_mode`.
+
0.10.0
......
diff --git a/providers/common/ai/docs/code_mode.rst
b/providers/common/ai/docs/code_mode.rst
index 90069cfd225..c3abf0a500b 100644
--- a/providers/common/ai/docs/code_mode.rst
+++ b/providers/common/ai/docs/code_mode.rst
@@ -25,9 +25,10 @@ Code mode
Experimental: this can change or be removed in a minor release of this
provider.
See :ref:`howto/stability`.
-Set ``code_mode=True`` to collapse the agent's tools into a single ``run_code``
-tool powered by the `Monty <https://github.com/pydantic/monty>`__ sandbox (via
-pydantic-ai-harness). Instead of one model round-trip per tool call, the model
+Pass pydantic-ai-harness's ``CodeMode`` capability in ``capabilities=`` (see
+:ref:`capabilities`) to collapse the
+agent's tools into a single ``run_code`` tool powered by the
+`Monty <https://github.com/pydantic/monty>`__ sandbox. Instead of one model
round-trip per tool call, the model
writes a single Python snippet that calls the tools as functions -- with loops,
conditionals, and ``asyncio.gather`` -- in one turn. For multi-tool workflows
this cuts round-trips and token use.
@@ -68,7 +69,7 @@ mode:
genuinely required rather than by default. See :doc:`sandbox/index` for the
backends and their limitations.
-The two are not exclusive: ``code_mode=True`` and a ``SandboxToolset`` can be
+The two are not exclusive: ``CodeMode`` and a ``SandboxToolset`` can be
enabled together, and the file tools fold into ``run_code`` while
``run_command``
stays a tool of its own.
@@ -76,16 +77,31 @@ Requires the ``code-mode`` extra::
pip install "apache-airflow-providers-common-ai[code-mode]"
+and the import ``from pydantic_ai_harness import CodeMode`` in the Dag file.
+
.. exampleinclude::
/../../ai/src/airflow/providers/common/ai/example_dags/example_agent.py
:language: python
:start-after: [START howto_operator_agent_code_mode]
:end-before: [END howto_operator_agent_code_mode]
-``code_mode=True`` is the same as adding ``CodeMode()`` at the end of
``capabilities=`` (see
-:ref:`capabilities`), except that the operator builds the capability when the
task runs, so the
-Dag file does not import ``pydantic_ai_harness``.
+``CodeMode`` takes its own arguments; see the
+`pydantic-ai-harness code mode docs
<https://ai.pydantic.dev/harness/code-mode/>`__. The one
+you are most likely to need is ``max_tool_calls``: each ``run_code`` snippet
may make at most
+that many nested tool calls, 100 by default. A snippet that asks for more
fails, the calls it
+already made are not undone, and the model has to split the work across
several snippets. If the
+agent fans out over a list, for example one tool call per row or per file, set
it above the
+largest list you expect.
+
+Code mode cannot be combined with ``durable=True`` (see :doc:`capabilities`
for which
+capabilities durable replay covers), and it turns off :doc:`tool_approval`. A
tool that needs
+approval and is called from ``run_code`` does not run: the model gets an error
back and may
+carry on without it, so the task can still succeed. To keep such a tool out of
``run_code``,
+list the other tools in ``CodeMode(tools=[...])``; called directly, it fails
the task.
+
+Importing ``CodeMode`` in the Dag file loads pydantic-ai-harness and Monty
each time the file is
+parsed. Keep code mode agents in their own Dag file if the rest of the file
does not need them.
.. note::
- Monty is pre-1.0. The ``code-mode`` extra is opt-in so its dependency churn
+ pydantic-ai-harness is pre-1.0. The ``code-mode`` extra is opt-in so its
dependency churn
never affects the base provider install.
diff --git a/providers/common/ai/docs/features.rst
b/providers/common/ai/docs/features.rst
index 3664b80fb84..b7df68cf6fe 100644
--- a/providers/common/ai/docs/features.rst
+++ b/providers/common/ai/docs/features.rst
@@ -28,8 +28,8 @@ you use. Each is a parameter on the operator or decorator.
- :doc:`message_history`: ``message_history`` carries a conversation across
agent runs.
- :doc:`capabilities`: ``capabilities=`` adds pydantic-ai capabilities such as
``Thinking`` and
``WebSearch``, and ``pydantic-ai-shields`` guardrails, to an agent.
-- :doc:`code_mode`: ``code_mode=True`` lets the model call several tools from
one Python
- snippet instead of one round trip per call.
+- :doc:`code_mode`: the ``CodeMode`` capability lets the model call several
tools from one
+ Python snippet instead of one round trip per call.
- :doc:`approval_gates`: ``require_approval=True`` pauses an LLM operator
until a person
approves, edits or rejects the output.
- :doc:`hitl_review`: ``enable_hitl_review=True`` opens an iterative review
loop on an agent,
diff --git a/providers/common/ai/docs/index.rst
b/providers/common/ai/docs/index.rst
index d6a32a58176..152e448f69a 100644
--- a/providers/common/ai/docs/index.rst
+++ b/providers/common/ai/docs/index.rst
@@ -227,7 +227,7 @@ Extra Dependencies
``mcp`` ``pydantic-ai-slim[mcp]>=2.33.0``
``modal`` ``apache-airflow-providers-modal``, ``modal>=1.5.2``
``opensandbox`` ``opensandbox>=1.1.0``
-``code-mode`` ``pydantic-ai-harness[codemode]>=0.3.0``
+``code-mode`` ``pydantic-ai-harness[codemode]>=0.24.0``
``shields`` ``pydantic-ai-shields>=0.3.4``
``skills`` ``apache-airflow-providers-git>=0.4.0``,
``pydantic-ai-skills>=1.2.0``
``avro`` ``fastavro>=1.10.0; python_version < "3.14"``,
``fastavro>=1.12.1; python_version >= "3.14"``
diff --git a/providers/common/ai/docs/operators/agent.rst
b/providers/common/ai/docs/operators/agent.rst
index 4144c104bbc..1f03bf00d0b 100644
--- a/providers/common/ai/docs/operators/agent.rst
+++ b/providers/common/ai/docs/operators/agent.rst
@@ -251,8 +251,8 @@ Five features have pages of their own:
on retry instead of paying for them again.
- :doc:`../capabilities`: pass pydantic-ai capabilities and
``pydantic-ai-shields`` guardrails
with ``capabilities=``.
-- :doc:`../code_mode`: set ``code_mode=True`` to collapse the agent's tools
into a single
- ``run_code`` tool the model drives by writing Python.
+- :doc:`../code_mode`: pass the ``CodeMode`` capability to collapse the
agent's tools into a
+ single ``run_code`` tool the model drives by writing Python.
- :doc:`../tool_approval`: mark tools that need a person's approval, and the
task pauses before
a marked call runs.
@@ -435,9 +435,6 @@ Parameters
when the steps it needs are cached. Clearing a failed task instance starts
a fresh budget but keeps the durable cache its attempts left behind, so what
the rerun replays from that cache is free there too.
-- ``code_mode``: When ``True``, wraps the agent's tools in a single
``run_code``
- tool that the model drives by writing Python, executed in the Monty sandbox.
- Requires the ``code-mode`` extra. Default ``False``. See :ref:`code-mode`.
- ``cache_prompt``: Ask the provider to cache the tool definitions, system
prompt and
conversation so later requests read them back at a discount. Default
``True``; a no-op for
providers that cache on their own. See :ref:`agent-prompt-caching`.
diff --git a/providers/common/ai/docs/sandbox/index.rst
b/providers/common/ai/docs/sandbox/index.rst
index 9022e591f7b..38fce059808 100644
--- a/providers/common/ai/docs/sandbox/index.rst
+++ b/providers/common/ai/docs/sandbox/index.rst
@@ -359,7 +359,7 @@ somewhere to work". This page has the worked scenarios
above and
the limitations below to read before designing a Dag around it.
Before reaching for it, check whether the actual need is narrower than that.
-``code_mode=True`` is a flag on ``AgentOperator``. It changes how the
+:ref:`Code mode <code-mode>` is a capability on ``AgentOperator``. It changes
how the
model invokes the tools it already has, letting it write code that calls
several
of them instead of emitting one call per step. It does not give the agent
somewhere
to run arbitrary code of its own. Because it needs no backend, it avoids the
diff --git a/providers/common/ai/docs/stability.rst
b/providers/common/ai/docs/stability.rst
index 2101c651253..8f3682b60b4 100644
--- a/providers/common/ai/docs/stability.rst
+++ b/providers/common/ai/docs/stability.rst
@@ -51,7 +51,7 @@ Pydantic AI toolsets, but the Pydantic AI class they inherit
from can change.
* - ``@task.agent`` and
:class:`~airflow.providers.common.ai.operators.agent.AgentOperator`
- Runs an agent with the model from ``llm_conn_id`` and the given
toolsets, and
returns its output under the same rules as ``@task.llm``. ``durable``,
- ``code_mode`` and per-tool approval are experimental; see below.
+ code mode and per-tool approval are experimental; see below.
* - ``message_history`` on ``AgentOperator``
- Seeds the run with the given conversation, as a list of messages or its
JSON form,
and publishes the finished conversation under the ``message_history``
XCom key so a
@@ -131,7 +131,7 @@ Everything this provider ships that is not in the table
above is experimental.
``on_tool_approval_timeout`` and ``tool_approval_assigned_users``;
:doc:`tool_approval`)
- New, and needs Airflow 3.3. How a paused run resumes may change.
- * - ``code_mode`` (:doc:`code_mode`), the Agent Skills toolset
+ * - Code mode (:doc:`code_mode`), the Agent Skills toolset
(:doc:`toolsets/skills`) and the ``shields`` extra (used in
:doc:`capabilities`)
- Thin integrations of packages outside this provider whose APIs are still
changing: ``pydantic-ai-harness``, ``pydantic-ai-skills`` and
diff --git a/providers/common/ai/docs/tool_approval.rst
b/providers/common/ai/docs/tool_approval.rst
index c568bd00a4c..8c1c9ebdbfb 100644
--- a/providers/common/ai/docs/tool_approval.rst
+++ b/providers/common/ai/docs/tool_approval.rst
@@ -127,11 +127,13 @@ Requirements and limits
- Airflow 3.3 or later. On older versions a tool marked for approval fails the
task, as it did before.
-- Not together with ``durable=True``, ``enable_hitl_review=True``, code mode
- (``code_mode=True`` or a ``CodeMode`` capability), or a ``SandboxToolset``
that
+- Not together with ``durable=True``, ``enable_hitl_review=True``, a
``CodeMode``
+ capability, or a ``SandboxToolset`` that
provisions its own sandbox. Each
assumes the run finishes in one go; that sandbox, for one, is destroyed when
the
- run pauses. With any of them, a marked tool fails the task. A
``SandboxToolset``
+ run pauses. With any of them, a marked tool fails the task, except under
code mode,
+ where a marked tool called from ``run_code`` does not run and the model may
carry on
+ without it (see :doc:`code_mode`). A ``SandboxToolset``
attached to a sandbox another task owns keeps its files through the pause,
so it
is allowed; the wait spends that sandbox's lifetime
(:ref:`sandbox-attach`).
diff --git a/providers/common/ai/docs/troubleshooting.rst
b/providers/common/ai/docs/troubleshooting.rst
index 9a0694fbc6b..a8c1c3978f1 100644
--- a/providers/common/ai/docs/troubleshooting.rst
+++ b/providers/common/ai/docs/troubleshooting.rst
@@ -44,8 +44,9 @@ Model and connection errors
version suffixes, for example), use the matching vendor connection type
instead of
the generic one.
-An ``ImportError`` for ``pydantic_ai.models.<vendor>`` or the vendor SDK
- The provider is installed without the extra for that vendor. Install it,
quoting the
+An ``ImportError`` for ``pydantic_ai.models.<vendor>``, the vendor SDK, or
``pydantic_ai_harness``
+ The provider is installed without the extra for that vendor, or without
``code-mode``
+ for ``pydantic_ai_harness``. Install it, quoting the
package name so the brackets survive the shell:
.. code-block:: bash
@@ -102,7 +103,7 @@ as a task failure.
or later. Upgrade Airflow, or use ``on_uncertain="fail"`` and drop the
review flags
on an older Airflow version. See :doc:`approval_gates` and
:doc:`hitl_review`.
-``durable=True and enable_hitl_review=True cannot be used together`` /
``durable=True and code_mode=True cannot be used together``
+``durable=True and enable_hitl_review=True cannot be used together`` /
``durable=True cannot be used with a CodeMode capability``
Durable replay assumes a stable step order across attempts, which neither
a human
review loop nor code mode provides. Pick one. See :doc:`durable_execution`.
@@ -110,9 +111,6 @@ as a task failure.
The post-review transcript is not recoverable today, so the operator
refuses rather
than silently dropping the reviewed turns. See :doc:`message_history`.
-``code_mode=True requires the 'code-mode' extra``
- Install ``apache-airflow-providers-common-ai[code-mode]``. See
:doc:`code_mode`.
-
``... does not support decision_policy yet``
Only ``LLMOperator`` and ``LLMBranchOperator`` honor a ``DecisionPolicy``
with a
confidence bar. The SQL, schema-compare and file-analysis operators run
their own
diff --git a/providers/common/ai/pyproject.toml
b/providers/common/ai/pyproject.toml
index c1663c2a81b..a14cd271839 100644
--- a/providers/common/ai/pyproject.toml
+++ b/providers/common/ai/pyproject.toml
@@ -112,9 +112,10 @@ dependencies = [
"opensandbox" = ["opensandbox>=1.1.0"]
# Code mode: collapse tool calls into a single `run_code` tool that the model
# drives by writing Python, executed in the Monty sandbox (pydantic-monty).
-# Enables AgentOperator(code_mode=True). Monty is pre-1.0; pinned here as an
-# opt-in extra so its churn never breaks the base provider install.
-"code-mode" = ["pydantic-ai-harness[codemode]>=0.3.0"]
+# Provides the CodeMode capability for
AgentOperator(capabilities=[CodeMode()]).
+# pydantic-ai-harness is pre-1.0; an opt-in extra so its churn never breaks the
+# base provider install.
+"code-mode" = ["pydantic-ai-harness[codemode]>=0.24.0"]
# Shield capabilities such as InputGuard, OutputGuard, ToolGuard, and
CostTracking.
"shields" = ["pydantic-ai-shields>=0.3.4"]
# Agent Skills (agentskills.io) support. pydantic-ai-skills provides the
toolset;
diff --git
a/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_agent.py
b/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_agent.py
index 900b44762d6..2eb36275f53 100644
---
a/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_agent.py
+++
b/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_agent.py
@@ -31,6 +31,11 @@ try:
except Exception:
SQLToolset = None # type: ignore[assignment,misc]
+try:
+ from pydantic_ai_harness import CodeMode
+except ImportError:
+ CodeMode = None # type: ignore[assignment,misc]
+
# [START howto_decorator_agent_structured_output_class]
# Pydantic output classes must be defined at module scope so downstream
@@ -276,7 +281,7 @@ example_agent_operator_hitl_review()
# [START howto_operator_agent_code_mode]
-if SQLToolset is not None:
+if SQLToolset is not None and CodeMode is not None:
@dag(tags=["example"])
def example_agent_operator_code_mode():
@@ -288,7 +293,7 @@ if SQLToolset is not None:
toolsets=[SQLToolset(db_conn_id="postgres_default",
allowed_tables=["customers", "orders"])],
# Requires the `code-mode` extra:
# pip install "apache-airflow-providers-common-ai[code-mode]"
- code_mode=True,
+ capabilities=[CodeMode(max_tool_calls=200)],
)
# [END howto_operator_agent_code_mode]
diff --git
a/providers/common/ai/src/airflow/providers/common/ai/operators/agent.py
b/providers/common/ai/src/airflow/providers/common/ai/operators/agent.py
index 6c4bccabf3c..26cc48d78e5 100644
--- a/providers/common/ai/src/airflow/providers/common/ai/operators/agent.py
+++ b/providers/common/ai/src/airflow/providers/common/ai/operators/agent.py
@@ -202,31 +202,6 @@ def _declares_agent_template_fields(toolset: Any) -> bool:
)
-def _build_code_mode() -> Any:
- """
- Return a pydantic-ai-harness ``CodeMode`` capability, or raise if not
installed.
-
- Kept here (not a module-level import) because ``pydantic-ai-harness`` is an
- optional dependency behind the ``code-mode`` extra; importing it eagerly
- would break installs that don't enable the extra.
- """
- try:
- from pydantic_ai_harness import CodeMode
- except ImportError as e:
- # Only report "extra not installed" when pydantic-ai-harness itself is
- # missing. A failure deeper in its import chain (a broken or missing
- # transitive dependency) is a different problem -- re-raise it as-is so
- # the real error isn't masked by a misleading "install the extra"
message.
- missing = e.name or ""
- if missing == "pydantic_ai_harness" or
missing.startswith("pydantic_ai_harness."):
- raise AirflowOptionalProviderFeatureException(
- "code_mode=True requires the 'code-mode' extra. Install it
with "
- '`pip install
"apache-airflow-providers-common-ai[code-mode]"`.'
- ) from e
- raise
- return CodeMode()
-
-
# CancellableAgentRunMixin must precede BaseOperator so its on_kill overrides
BaseOperator's
# no-op. The other mixins only add methods, so they can trail BaseOperator.
See the MRO guard
# test in tests/unit/common/ai/mixins/test_cancellable_run.py.
@@ -355,20 +330,8 @@ class AgentOperator(CancellableAgentRunMixin,
BaseOperator, HITLReviewMixin):
Cannot be combined with a ``SandboxToolset`` (raises), attached or
not: a replayed tool result describes a workspace state the replay did
not reproduce, and the first call that misses the cache runs against
- whatever the sandbox holds now.
- :param code_mode: Experimental. When ``True``, wraps the agent's tools in
a single
- ``run_code`` tool powered by the Monty sandbox (pydantic-ai-harness
- ``CodeMode``). Instead of one model round-trip per tool call, the model
- writes Python that calls the tools as functions, with loops and
- ``asyncio.gather``, in one turn. The generated code runs in Monty's
- deny-by-default sandbox; the tools it calls still run in the worker, so
- ``code_mode`` does not widen what the tools can reach -- it only
changes
- how the model invokes them. Requires the ``code-mode`` extra
- (``pip install "apache-airflow-providers-common-ai[code-mode]"``).
- Cannot be combined with ``durable=True`` (durable replay assumes a
- stable per-step call order that code mode does not guarantee), whether
- code mode comes from this flag or from a ``CodeMode`` capability.
- Default ``False``.
+ whatever the sandbox holds now. Cannot be combined with a
pydantic-ai-harness
+ ``CodeMode`` capability (raises).
:param cache_prompt: When ``True`` (default), asks the provider to cache
the
tool definitions, system prompt and conversation so far, so the next
request in the run -- and a mapped task's other instances within the
@@ -435,10 +398,11 @@ class AgentOperator(CancellableAgentRunMixin,
BaseOperator, HITLReviewMixin):
(with the reviewer's reason, when given) and carries on without it. A task
instance asks at most once per Dag run, across retries and clears; a second
request fails the task. ``usage_limits`` applies to both sides of the
pause.
- Not available together with ``durable``, ``enable_hitl_review``, code mode
- (``code_mode=True`` or a ``CodeMode`` capability), or a ``SandboxToolset``
+ Not available together with ``durable``, ``enable_hitl_review``, a
``CodeMode``
+ capability, or a ``SandboxToolset``
that provisions its own sandbox; there, a tool that requires approval fails
- the task as before. A ``SandboxToolset`` attached to a
+ the task as before, except one called from inside ``CodeMode``'s
``run_code``,
+ which does not run and is reported back to the model. A ``SandboxToolset``
attached to a
sandbox another task owns is fine: the sandbox outlives the pause.
:param tool_approval_timeout: Experimental. How long the pause waits for a
decision.
@@ -493,7 +457,6 @@ class AgentOperator(CancellableAgentRunMixin, BaseOperator,
HITLReviewMixin):
agent_params: dict[str, Any] | None = None,
usage_limits: UsageLimits | dict[str, Any] | None = None,
durable: bool = False,
- code_mode: bool = False,
cache_prompt: bool = True,
message_history: list[ModelMessage] | str | bytes | None = None,
# Agent feedback parameters
@@ -528,7 +491,6 @@ class AgentOperator(CancellableAgentRunMixin, BaseOperator,
HITLReviewMixin):
self.message_history = message_history
self.durable = durable
- self.code_mode = code_mode
self.cache_prompt = cache_prompt
# Populated per run in ``execute`` when durable=True. Declared here so
@@ -556,22 +518,14 @@ class AgentOperator(CancellableAgentRunMixin,
BaseOperator, HITLReviewMixin):
if durable and enable_hitl_review:
raise ValueError("durable=True and enable_hitl_review=True cannot
be used together.")
- if durable and code_mode:
+ if durable and _contains_code_mode(self._declared_capabilities):
# Durable replay caches individual model/tool steps via
CachingModel /
# CachingToolset and a shared step counter that assumes a stable
call
# order across runs. Code mode collapses tools into one
``run_code``
# tool and lets the model emit arbitrary Python, so step counts and
# ordering can differ between the original run and a retry,
breaking
# replay. Reject the combination rather than silently
mis-replaying.
- raise ValueError("durable=True and code_mode=True cannot be used
together.")
-
- if (durable or code_mode) and
_contains_code_mode(self._declared_capabilities):
- if durable:
- # The same conflict as code_mode=True, reached through the
capability itself.
- raise ValueError("durable=True cannot be used with a CodeMode
capability.")
- # code_mode=True adds a second CodeMode, and pydantic-ai then
fails the run on a
- # duplicate ``run_code`` tool without saying where the second one
came from.
- raise ValueError("code_mode=True adds a CodeMode capability; pass
one or the other, not both.")
+ raise ValueError("durable=True cannot be used with a CodeMode
capability.")
if message_history is not None and enable_hitl_review:
# The post-review transcript is not recoverable today
(run_hitl_review
@@ -768,8 +722,6 @@ class AgentOperator(CancellableAgentRunMixin, BaseOperator,
HITLReviewMixin):
# ``toolsets=`` wrapping above, so their results would re-execute
on
# every retry instead of replaying; wrap their inner toolset too.
capabilities = self._build_durable_capabilities(capabilities,
storage, counter)
- if self.code_mode:
- capabilities.append(_build_code_mode())
if self.cache_prompt:
capabilities.append(PromptCaching())
if capabilities:
@@ -794,7 +746,6 @@ class AgentOperator(CancellableAgentRunMixin, BaseOperator,
HITLReviewMixin):
not AIRFLOW_V_3_3_PLUS
or self.durable
or self.enable_hitl_review
- or self.code_mode
or _contains_code_mode(self._declared_capabilities)
):
return False
@@ -1229,7 +1180,7 @@ class AgentOperator(CancellableAgentRunMixin,
BaseOperator, HITLReviewMixin):
raise UnsupportedToolDeferralError(
f"The agent called tools that need approval ({pending_names}),
but tool approval "
"needs Airflow 3.3+ and is not available with durable,
enable_hitl_review, "
- "code mode (code_mode=True or a CodeMode capability) or a
SandboxToolset."
+ "a CodeMode capability or a SandboxToolset."
)
store = context["task_state_store"]
if store.get(_TOOL_APPROVAL_REQUESTED_KEY):
diff --git a/providers/common/ai/tests/unit/common/ai/operators/test_agent.py
b/providers/common/ai/tests/unit/common/ai/operators/test_agent.py
index 3823eacee0e..01273b99d8a 100644
--- a/providers/common/ai/tests/unit/common/ai/operators/test_agent.py
+++ b/providers/common/ai/tests/unit/common/ai/operators/test_agent.py
@@ -64,7 +64,7 @@ from airflow.providers.common.ai.durable.base import (
from airflow.providers.common.ai.durable.caching_toolset import CachingToolset
from airflow.providers.common.ai.durable.step_counter import DurableStepCounter
from airflow.providers.common.ai.durable.storage import DurableStorage
-from airflow.providers.common.ai.operators.agent import AgentOperator,
HITLReviewLink, _build_code_mode
+from airflow.providers.common.ai.operators.agent import AgentOperator,
HITLReviewLink
from airflow.providers.common.ai.sandbox.base import (
HOLDER_TAG,
OWNER_TAG,
@@ -860,8 +860,8 @@ class TestAgentOperatorExecute:
assert create_call[1]["model_settings"] == {"temperature": 0}
@patch("airflow.providers.common.ai.operators.agent.PydanticAIHook",
autospec=True)
- def test_code_mode_default_off_no_capabilities(self, mock_hook_cls,
make_mock_run_result):
- """code_mode defaults to False, so with cache_prompt off no
capabilities are injected."""
+ def test_no_capabilities_injected_with_cache_prompt_off(self,
mock_hook_cls, make_mock_run_result):
+ """With cache_prompt off and no capabilities passed, create_agent gets
no capabilities."""
mock_hook_cls.get_hook.return_value.create_agent.return_value =
_make_mock_agent(
"ok", make_mock_run_result
)
@@ -878,83 +878,6 @@ class TestAgentOperatorExecute:
create_call =
mock_hook_cls.get_hook.return_value.create_agent.call_args
assert "capabilities" not in create_call[1]
- @patch("airflow.providers.common.ai.operators.agent._build_code_mode",
return_value="CM")
- @patch("airflow.providers.common.ai.operators.agent.PydanticAIHook",
autospec=True)
- def test_code_mode_injects_capability(self, mock_hook_cls, mock_build,
make_mock_run_result):
- """code_mode=True appends a CodeMode capability passed to
create_agent."""
- mock_hook_cls.get_hook.return_value.create_agent.return_value =
_make_mock_agent(
- "ok", make_mock_run_result
- )
-
- op = AgentOperator(
- task_id="t",
- prompt="hi",
- llm_conn_id="my_llm",
- toolsets=[MagicMock(spec=AbstractToolset)],
- code_mode=True,
- cache_prompt=False,
- )
- op.execute(context=_make_context())
-
- create_call =
mock_hook_cls.get_hook.return_value.create_agent.call_args
- assert create_call[1]["capabilities"] == ["CM"]
- mock_build.assert_called_once()
-
- @patch("airflow.providers.common.ai.operators.agent._build_code_mode",
return_value="CM")
- @patch("airflow.providers.common.ai.operators.agent.PydanticAIHook",
autospec=True)
- def test_code_mode_appends_to_existing_capabilities(
- self, mock_hook_cls, mock_build, make_mock_run_result
- ):
- """A user-supplied capability via agent_params is preserved alongside
CodeMode."""
- mock_hook_cls.get_hook.return_value.create_agent.return_value =
_make_mock_agent(
- "ok", make_mock_run_result
- )
-
- op = AgentOperator(
- task_id="t",
- prompt="hi",
- llm_conn_id="my_llm",
- code_mode=True,
- cache_prompt=False,
- agent_params={"capabilities": ["existing"]},
- )
- op.execute(context=_make_context())
-
- create_call =
mock_hook_cls.get_hook.return_value.create_agent.call_args
- assert create_call[1]["capabilities"] == ["existing", "CM"]
-
- def test_build_code_mode_missing_harness_raises(self):
- """_build_code_mode raises the optional-feature error when harness is
absent."""
- with patch.dict(sys.modules, {"pydantic_ai_harness": None}):
- with pytest.raises(AirflowOptionalProviderFeatureException,
match="code-mode"):
- _build_code_mode()
-
- def test_build_code_mode_reraises_unrelated_import_error(self):
- """A broken transitive import inside the harness is re-raised, not
masked as 'extra missing'."""
- real_import = __import__
-
- def fake_import(name, *args, **kwargs):
- if name == "pydantic_ai_harness":
- raise ModuleNotFoundError("No module named 'a_broken_dep'",
name="a_broken_dep")
- return real_import(name, *args, **kwargs)
-
- with patch("builtins.__import__", side_effect=fake_import):
- with pytest.raises(ModuleNotFoundError, match="a_broken_dep"):
- _build_code_mode()
-
- @patch("airflow.providers.common.ai.operators.agent._build_code_mode")
- def test_code_mode_not_built_at_init(self, mock_build):
- """code_mode is serialization-safe: the CodeMode capability is built
lazily in
- _build_agent, never at construction time (so nothing non-serializable
is stored)."""
- op = AgentOperator(task_id="t", prompt="hi", llm_conn_id="my_llm",
code_mode=True)
- mock_build.assert_not_called()
- assert op.code_mode is True
-
- def test_durable_and_code_mode_rejected(self):
- """durable and code_mode cannot be combined (durable replay assumes
stable step order)."""
- with pytest.raises(ValueError, match="durable=True and
code_mode=True"):
- AgentOperator(task_id="t", prompt="hi", llm_conn_id="my_llm",
durable=True, code_mode=True)
-
@patch("airflow.providers.common.ai.operators.agent.PydanticAIHook",
autospec=True)
def test_cache_prompt_default_on_appends_capability_last(self,
mock_hook_cls, make_mock_run_result):
"""cache_prompt defaults to True and adds PromptCaching after any user
capability."""
@@ -1366,12 +1289,6 @@ class TestAgentOperatorCapabilities:
assert op._supports_tool_approval() is True
- def
test_code_mode_flag_and_code_mode_capability_are_refused_together(self,
fake_harness):
- with pytest.raises(ValueError, match="one or the other"):
- AgentOperator(
- task_id="t", prompt="p", llm_conn_id="llm", code_mode=True,
capabilities=[_FakeCodeMode()]
- )
-
def test_capability_function_is_left_for_the_run_to_resolve(self,
fake_harness):
def build(ctx):
return Thinking()
diff --git
a/providers/common/ai/tests/unit/common/ai/operators/test_agent_tool_approval.py
b/providers/common/ai/tests/unit/common/ai/operators/test_agent_tool_approval.py
index 09e2fe93732..7334b98fc2e 100644
---
a/providers/common/ai/tests/unit/common/ai/operators/test_agent_tool_approval.py
+++
b/providers/common/ai/tests/unit/common/ai/operators/test_agent_tool_approval.py
@@ -551,7 +551,6 @@ class TestWhenApprovalApplies:
"kwargs",
[
pytest.param({"durable": True}, id="durable"),
- pytest.param({"code_mode": True}, id="code_mode"),
pytest.param({"enable_hitl_review": True}, id="hitl_review"),
pytest.param({"toolsets": [SandboxToolset(_NoopBackend())]},
id="sandbox"),
pytest.param(
diff --git a/uv.lock b/uv.lock
index f4bc231f9d3..ce3eafa7ab1 100644
--- a/uv.lock
+++ b/uv.lock
@@ -4763,7 +4763,7 @@ requires-dist = [
{ name = "opensandbox", marker = "extra == 'opensandbox'", specifier =
">=1.1.0" },
{ name = "pyarrow", marker = "python_full_version >= '3.14' and extra ==
'parquet'", specifier = ">=22.0.0" },
{ name = "pyarrow", marker = "python_full_version < '3.14' and extra ==
'parquet'", specifier = ">=18.0.0" },
- { name = "pydantic-ai-harness", extras = ["codemode"], marker = "extra ==
'code-mode'", specifier = ">=0.3.0" },
+ { name = "pydantic-ai-harness", extras = ["codemode"], marker = "extra ==
'code-mode'", specifier = ">=0.24.0" },
{ name = "pydantic-ai-shields", marker = "extra == 'shields'", specifier =
">=0.3.4" },
{ name = "pydantic-ai-skills", marker = "extra == 'skills'", specifier =
">=1.2.0" },
{ name = "pydantic-ai-slim", specifier = ">=2.33.0" },