This is an automated email from the ASF dual-hosted git repository.

kaxil pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/airflow.git


The following commit(s) were added to refs/heads/main by this push:
     new ac147d07ffe Remove `code_mode` from `AgentOperator` in favor of the 
`CodeMode` capability (#74312)
ac147d07ffe is described below

commit ac147d07ffe3f1d390a2fbc062be80000d39e480
Author: Kaxil Naik <[email protected]>
AuthorDate: Tue Oct 6 08:28:16 2026 +0100

    Remove `code_mode` from `AgentOperator` in favor of the `CodeMode` 
capability (#74312)
    
    code_mode=True only appended a bare CodeMode() to the agent's capabilities,
    which capabilities= now does directly. The flag hid CodeMode's own 
arguments,
    most importantly max_tool_calls, whose default of 100 nested calls per
    run_code snippet is too low for agents that fan out over a list.
---
 providers/common/ai/docs/capabilities.rst          |  6 +-
 providers/common/ai/docs/changelog.rst             | 10 +++
 providers/common/ai/docs/code_mode.rst             | 32 ++++++--
 providers/common/ai/docs/features.rst              |  4 +-
 providers/common/ai/docs/index.rst                 |  2 +-
 providers/common/ai/docs/operators/agent.rst       |  7 +-
 providers/common/ai/docs/sandbox/index.rst         |  2 +-
 providers/common/ai/docs/stability.rst             |  4 +-
 providers/common/ai/docs/tool_approval.rst         |  8 +-
 providers/common/ai/docs/troubleshooting.rst       | 10 +--
 providers/common/ai/pyproject.toml                 |  7 +-
 .../common/ai/example_dags/example_agent.py        |  9 ++-
 .../airflow/providers/common/ai/operators/agent.py | 67 +++-------------
 .../tests/unit/common/ai/operators/test_agent.py   | 89 +---------------------
 .../ai/operators/test_agent_tool_approval.py       |  1 -
 uv.lock                                            |  2 +-
 16 files changed, 78 insertions(+), 182 deletions(-)

diff --git a/providers/common/ai/docs/capabilities.rst 
b/providers/common/ai/docs/capabilities.rst
index eeca7d716fd..0284396dff6 100644
--- a/providers/common/ai/docs/capabilities.rst
+++ b/providers/common/ai/docs/capabilities.rst
@@ -112,9 +112,9 @@ again. Whether a capability's work is replayed depends on 
where it runs:
         and capabilities from an agent spec file
       - Tools run again. Pass tools you need replayed in ``toolsets=`` instead.
     * - pydantic-ai-harness ``CodeMode``
-      - Not allowed: the operator raises ``ValueError``, as it does for 
``code_mode=True``. This
-        includes a ``CodeMode`` inside a ``CombinedCapability`` or a wrapper 
such as
-        ``PrefixTools``, but not one a capability function builds when the run 
starts.
+      - Not allowed: the operator raises ``ValueError``. This includes a 
``CodeMode`` inside a
+        ``CombinedCapability`` or a wrapper such as ``PrefixTools``, but not 
one a capability
+        function builds when the run starts.
 
 See :doc:`durable_execution` for how the cache works.
 
diff --git a/providers/common/ai/docs/changelog.rst 
b/providers/common/ai/docs/changelog.rst
index 96b4ea26f35..330b23aa439 100644
--- a/providers/common/ai/docs/changelog.rst
+++ b/providers/common/ai/docs/changelog.rst
@@ -51,6 +51,16 @@ Changelog
   result or ``result["sample_columns"]`` on a summary. A summary carries no 
``columns`` key, so
   ``result["columns"]`` raises ``KeyError`` on any table wide enough to be 
summarized.
 
+.. note::
+  ``AgentOperator`` no longer accepts ``code_mode``. Pass the 
pydantic-ai-harness capability
+  instead: replace ``code_mode=True`` with ``capabilities=[CodeMode()]``, using
+  ``from pydantic_ai_harness import CodeMode``. The capability takes its own 
arguments, such
+  as ``max_tool_calls``. If the task also passes 
``agent_params={"capabilities": [...]}``,
+  move those into ``capabilities=`` too, since the two cannot be combined. The 
import runs
+  when the Dag file is parsed, so the ``code-mode`` extra is now needed by the 
Dag processor
+  as well as the workers. The extra's floor is now 
``pydantic-ai-harness>=0.24.0``, the first
+  release with ``max_tool_calls``. See :doc:`code_mode`.
+
 0.10.0
 ......
 
diff --git a/providers/common/ai/docs/code_mode.rst 
b/providers/common/ai/docs/code_mode.rst
index 90069cfd225..c3abf0a500b 100644
--- a/providers/common/ai/docs/code_mode.rst
+++ b/providers/common/ai/docs/code_mode.rst
@@ -25,9 +25,10 @@ Code mode
     Experimental: this can change or be removed in a minor release of this 
provider.
     See :ref:`howto/stability`.
 
-Set ``code_mode=True`` to collapse the agent's tools into a single ``run_code``
-tool powered by the `Monty <https://github.com/pydantic/monty>`__ sandbox (via
-pydantic-ai-harness). Instead of one model round-trip per tool call, the model
+Pass pydantic-ai-harness's ``CodeMode`` capability in ``capabilities=`` (see
+:ref:`capabilities`) to collapse the
+agent's tools into a single ``run_code`` tool powered by the
+`Monty <https://github.com/pydantic/monty>`__ sandbox. Instead of one model 
round-trip per tool call, the model
 writes a single Python snippet that calls the tools as functions -- with loops,
 conditionals, and ``asyncio.gather`` -- in one turn. For multi-tool workflows
 this cuts round-trips and token use.
@@ -68,7 +69,7 @@ mode:
   genuinely required rather than by default. See :doc:`sandbox/index` for the
   backends and their limitations.
 
-The two are not exclusive: ``code_mode=True`` and a ``SandboxToolset`` can be
+The two are not exclusive: ``CodeMode`` and a ``SandboxToolset`` can be
 enabled together, and the file tools fold into ``run_code`` while 
``run_command``
 stays a tool of its own.
 
@@ -76,16 +77,31 @@ Requires the ``code-mode`` extra::
 
     pip install "apache-airflow-providers-common-ai[code-mode]"
 
+and the import ``from pydantic_ai_harness import CodeMode`` in the Dag file.
+
 .. exampleinclude:: 
/../../ai/src/airflow/providers/common/ai/example_dags/example_agent.py
     :language: python
     :start-after: [START howto_operator_agent_code_mode]
     :end-before: [END howto_operator_agent_code_mode]
 
-``code_mode=True`` is the same as adding ``CodeMode()`` at the end of 
``capabilities=`` (see
-:ref:`capabilities`), except that the operator builds the capability when the 
task runs, so the
-Dag file does not import ``pydantic_ai_harness``.
+``CodeMode`` takes its own arguments; see the
+`pydantic-ai-harness code mode docs 
<https://ai.pydantic.dev/harness/code-mode/>`__. The one
+you are most likely to need is ``max_tool_calls``: each ``run_code`` snippet 
may make at most
+that many nested tool calls, 100 by default. A snippet that asks for more 
fails, the calls it
+already made are not undone, and the model has to split the work across 
several snippets. If the
+agent fans out over a list, for example one tool call per row or per file, set 
it above the
+largest list you expect.
+
+Code mode cannot be combined with ``durable=True`` (see :doc:`capabilities` 
for which
+capabilities durable replay covers), and it turns off :doc:`tool_approval`. A 
tool that needs
+approval and is called from ``run_code`` does not run: the model gets an error 
back and may
+carry on without it, so the task can still succeed. To keep such a tool out of 
``run_code``,
+list the other tools in ``CodeMode(tools=[...])``; called directly, it fails 
the task.
+
+Importing ``CodeMode`` in the Dag file loads pydantic-ai-harness and Monty 
each time the file is
+parsed. Keep code mode agents in their own Dag file if the rest of the file 
does not need them.
 
 .. note::
 
-    Monty is pre-1.0. The ``code-mode`` extra is opt-in so its dependency churn
+    pydantic-ai-harness is pre-1.0. The ``code-mode`` extra is opt-in so its 
dependency churn
     never affects the base provider install.
diff --git a/providers/common/ai/docs/features.rst 
b/providers/common/ai/docs/features.rst
index 3664b80fb84..b7df68cf6fe 100644
--- a/providers/common/ai/docs/features.rst
+++ b/providers/common/ai/docs/features.rst
@@ -28,8 +28,8 @@ you use. Each is a parameter on the operator or decorator.
 - :doc:`message_history`: ``message_history`` carries a conversation across 
agent runs.
 - :doc:`capabilities`: ``capabilities=`` adds pydantic-ai capabilities such as 
``Thinking`` and
   ``WebSearch``, and ``pydantic-ai-shields`` guardrails, to an agent.
-- :doc:`code_mode`: ``code_mode=True`` lets the model call several tools from 
one Python
-  snippet instead of one round trip per call.
+- :doc:`code_mode`: the ``CodeMode`` capability lets the model call several 
tools from one
+  Python snippet instead of one round trip per call.
 - :doc:`approval_gates`: ``require_approval=True`` pauses an LLM operator 
until a person
   approves, edits or rejects the output.
 - :doc:`hitl_review`: ``enable_hitl_review=True`` opens an iterative review 
loop on an agent,
diff --git a/providers/common/ai/docs/index.rst 
b/providers/common/ai/docs/index.rst
index d6a32a58176..152e448f69a 100644
--- a/providers/common/ai/docs/index.rst
+++ b/providers/common/ai/docs/index.rst
@@ -227,7 +227,7 @@ Extra            Dependencies
 ``mcp``          ``pydantic-ai-slim[mcp]>=2.33.0``
 ``modal``        ``apache-airflow-providers-modal``, ``modal>=1.5.2``
 ``opensandbox``  ``opensandbox>=1.1.0``
-``code-mode``    ``pydantic-ai-harness[codemode]>=0.3.0``
+``code-mode``    ``pydantic-ai-harness[codemode]>=0.24.0``
 ``shields``      ``pydantic-ai-shields>=0.3.4``
 ``skills``       ``apache-airflow-providers-git>=0.4.0``, 
``pydantic-ai-skills>=1.2.0``
 ``avro``         ``fastavro>=1.10.0; python_version < "3.14"``, 
``fastavro>=1.12.1; python_version >= "3.14"``
diff --git a/providers/common/ai/docs/operators/agent.rst 
b/providers/common/ai/docs/operators/agent.rst
index 4144c104bbc..1f03bf00d0b 100644
--- a/providers/common/ai/docs/operators/agent.rst
+++ b/providers/common/ai/docs/operators/agent.rst
@@ -251,8 +251,8 @@ Five features have pages of their own:
   on retry instead of paying for them again.
 - :doc:`../capabilities`: pass pydantic-ai capabilities and 
``pydantic-ai-shields`` guardrails
   with ``capabilities=``.
-- :doc:`../code_mode`: set ``code_mode=True`` to collapse the agent's tools 
into a single
-  ``run_code`` tool the model drives by writing Python.
+- :doc:`../code_mode`: pass the ``CodeMode`` capability to collapse the 
agent's tools into a
+  single ``run_code`` tool the model drives by writing Python.
 - :doc:`../tool_approval`: mark tools that need a person's approval, and the 
task pauses before
   a marked call runs.
 
@@ -435,9 +435,6 @@ Parameters
   when the steps it needs are cached. Clearing a failed task instance starts
   a fresh budget but keeps the durable cache its attempts left behind, so what
   the rerun replays from that cache is free there too.
-- ``code_mode``: When ``True``, wraps the agent's tools in a single 
``run_code``
-  tool that the model drives by writing Python, executed in the Monty sandbox.
-  Requires the ``code-mode`` extra. Default ``False``. See :ref:`code-mode`.
 - ``cache_prompt``: Ask the provider to cache the tool definitions, system 
prompt and
   conversation so later requests read them back at a discount. Default 
``True``; a no-op for
   providers that cache on their own. See :ref:`agent-prompt-caching`.
diff --git a/providers/common/ai/docs/sandbox/index.rst 
b/providers/common/ai/docs/sandbox/index.rst
index 9022e591f7b..38fce059808 100644
--- a/providers/common/ai/docs/sandbox/index.rst
+++ b/providers/common/ai/docs/sandbox/index.rst
@@ -359,7 +359,7 @@ somewhere to work". This page has the worked scenarios 
above and
 the limitations below to read before designing a Dag around it.
 
 Before reaching for it, check whether the actual need is narrower than that.
-``code_mode=True`` is a flag on ``AgentOperator``. It changes how the
+:ref:`Code mode <code-mode>` is a capability on ``AgentOperator``. It changes 
how the
 model invokes the tools it already has, letting it write code that calls 
several
 of them instead of emitting one call per step. It does not give the agent 
somewhere
 to run arbitrary code of its own. Because it needs no backend, it avoids the
diff --git a/providers/common/ai/docs/stability.rst 
b/providers/common/ai/docs/stability.rst
index 2101c651253..8f3682b60b4 100644
--- a/providers/common/ai/docs/stability.rst
+++ b/providers/common/ai/docs/stability.rst
@@ -51,7 +51,7 @@ Pydantic AI toolsets, but the Pydantic AI class they inherit 
from can change.
    * - ``@task.agent`` and 
:class:`~airflow.providers.common.ai.operators.agent.AgentOperator`
      - Runs an agent with the model from ``llm_conn_id`` and the given 
toolsets, and
        returns its output under the same rules as ``@task.llm``. ``durable``,
-       ``code_mode`` and per-tool approval are experimental; see below.
+       code mode and per-tool approval are experimental; see below.
    * - ``message_history`` on ``AgentOperator``
      - Seeds the run with the given conversation, as a list of messages or its 
JSON form,
        and publishes the finished conversation under the ``message_history`` 
XCom key so a
@@ -131,7 +131,7 @@ Everything this provider ships that is not in the table 
above is experimental.
        ``on_tool_approval_timeout`` and ``tool_approval_assigned_users``;
        :doc:`tool_approval`)
      - New, and needs Airflow 3.3. How a paused run resumes may change.
-   * - ``code_mode`` (:doc:`code_mode`), the Agent Skills toolset
+   * - Code mode (:doc:`code_mode`), the Agent Skills toolset
        (:doc:`toolsets/skills`) and the ``shields`` extra (used in 
:doc:`capabilities`)
      - Thin integrations of packages outside this provider whose APIs are still
        changing: ``pydantic-ai-harness``, ``pydantic-ai-skills`` and
diff --git a/providers/common/ai/docs/tool_approval.rst 
b/providers/common/ai/docs/tool_approval.rst
index c568bd00a4c..8c1c9ebdbfb 100644
--- a/providers/common/ai/docs/tool_approval.rst
+++ b/providers/common/ai/docs/tool_approval.rst
@@ -127,11 +127,13 @@ Requirements and limits
 
 - Airflow 3.3 or later. On older versions a tool marked for approval fails the
   task, as it did before.
-- Not together with ``durable=True``, ``enable_hitl_review=True``, code mode
-  (``code_mode=True`` or a ``CodeMode`` capability), or a ``SandboxToolset`` 
that
+- Not together with ``durable=True``, ``enable_hitl_review=True``, a 
``CodeMode``
+  capability, or a ``SandboxToolset`` that
   provisions its own sandbox. Each
   assumes the run finishes in one go; that sandbox, for one, is destroyed when 
the
-  run pauses. With any of them, a marked tool fails the task. A 
``SandboxToolset``
+  run pauses. With any of them, a marked tool fails the task, except under 
code mode,
+  where a marked tool called from ``run_code`` does not run and the model may 
carry on
+  without it (see :doc:`code_mode`). A ``SandboxToolset``
   attached to a sandbox another task owns keeps its files through the pause, 
so it
   is allowed; the wait spends that sandbox's lifetime
   (:ref:`sandbox-attach`).
diff --git a/providers/common/ai/docs/troubleshooting.rst 
b/providers/common/ai/docs/troubleshooting.rst
index 9a0694fbc6b..a8c1c3978f1 100644
--- a/providers/common/ai/docs/troubleshooting.rst
+++ b/providers/common/ai/docs/troubleshooting.rst
@@ -44,8 +44,9 @@ Model and connection errors
     version suffixes, for example), use the matching vendor connection type 
instead of
     the generic one.
 
-An ``ImportError`` for ``pydantic_ai.models.<vendor>`` or the vendor SDK
-    The provider is installed without the extra for that vendor. Install it, 
quoting the
+An ``ImportError`` for ``pydantic_ai.models.<vendor>``, the vendor SDK, or 
``pydantic_ai_harness``
+    The provider is installed without the extra for that vendor, or without 
``code-mode``
+    for ``pydantic_ai_harness``. Install it, quoting the
     package name so the brackets survive the shell:
 
     .. code-block:: bash
@@ -102,7 +103,7 @@ as a task failure.
     or later. Upgrade Airflow, or use ``on_uncertain="fail"`` and drop the 
review flags
     on an older Airflow version. See :doc:`approval_gates` and 
:doc:`hitl_review`.
 
-``durable=True and enable_hitl_review=True cannot be used together`` / 
``durable=True and code_mode=True cannot be used together``
+``durable=True and enable_hitl_review=True cannot be used together`` / 
``durable=True cannot be used with a CodeMode capability``
     Durable replay assumes a stable step order across attempts, which neither 
a human
     review loop nor code mode provides. Pick one. See :doc:`durable_execution`.
 
@@ -110,9 +111,6 @@ as a task failure.
     The post-review transcript is not recoverable today, so the operator 
refuses rather
     than silently dropping the reviewed turns. See :doc:`message_history`.
 
-``code_mode=True requires the 'code-mode' extra``
-    Install ``apache-airflow-providers-common-ai[code-mode]``. See 
:doc:`code_mode`.
-
 ``... does not support decision_policy yet``
     Only ``LLMOperator`` and ``LLMBranchOperator`` honor a ``DecisionPolicy`` 
with a
     confidence bar. The SQL, schema-compare and file-analysis operators run 
their own
diff --git a/providers/common/ai/pyproject.toml 
b/providers/common/ai/pyproject.toml
index c1663c2a81b..a14cd271839 100644
--- a/providers/common/ai/pyproject.toml
+++ b/providers/common/ai/pyproject.toml
@@ -112,9 +112,10 @@ dependencies = [
 "opensandbox" = ["opensandbox>=1.1.0"]
 # Code mode: collapse tool calls into a single `run_code` tool that the model
 # drives by writing Python, executed in the Monty sandbox (pydantic-monty).
-# Enables AgentOperator(code_mode=True). Monty is pre-1.0; pinned here as an
-# opt-in extra so its churn never breaks the base provider install.
-"code-mode" = ["pydantic-ai-harness[codemode]>=0.3.0"]
+# Provides the CodeMode capability for 
AgentOperator(capabilities=[CodeMode()]).
+# pydantic-ai-harness is pre-1.0; an opt-in extra so its churn never breaks the
+# base provider install.
+"code-mode" = ["pydantic-ai-harness[codemode]>=0.24.0"]
 # Shield capabilities such as InputGuard, OutputGuard, ToolGuard, and 
CostTracking.
 "shields" = ["pydantic-ai-shields>=0.3.4"]
 # Agent Skills (agentskills.io) support. pydantic-ai-skills provides the 
toolset;
diff --git 
a/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_agent.py
 
b/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_agent.py
index 900b44762d6..2eb36275f53 100644
--- 
a/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_agent.py
+++ 
b/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_agent.py
@@ -31,6 +31,11 @@ try:
 except Exception:
     SQLToolset = None  # type: ignore[assignment,misc]
 
+try:
+    from pydantic_ai_harness import CodeMode
+except ImportError:
+    CodeMode = None  # type: ignore[assignment,misc]
+
 
 # [START howto_decorator_agent_structured_output_class]
 # Pydantic output classes must be defined at module scope so downstream
@@ -276,7 +281,7 @@ example_agent_operator_hitl_review()
 
 
 # [START howto_operator_agent_code_mode]
-if SQLToolset is not None:
+if SQLToolset is not None and CodeMode is not None:
 
     @dag(tags=["example"])
     def example_agent_operator_code_mode():
@@ -288,7 +293,7 @@ if SQLToolset is not None:
             toolsets=[SQLToolset(db_conn_id="postgres_default", 
allowed_tables=["customers", "orders"])],
             # Requires the `code-mode` extra:
             #   pip install "apache-airflow-providers-common-ai[code-mode]"
-            code_mode=True,
+            capabilities=[CodeMode(max_tool_calls=200)],
         )
 
     # [END howto_operator_agent_code_mode]
diff --git 
a/providers/common/ai/src/airflow/providers/common/ai/operators/agent.py 
b/providers/common/ai/src/airflow/providers/common/ai/operators/agent.py
index 6c4bccabf3c..26cc48d78e5 100644
--- a/providers/common/ai/src/airflow/providers/common/ai/operators/agent.py
+++ b/providers/common/ai/src/airflow/providers/common/ai/operators/agent.py
@@ -202,31 +202,6 @@ def _declares_agent_template_fields(toolset: Any) -> bool:
     )
 
 
-def _build_code_mode() -> Any:
-    """
-    Return a pydantic-ai-harness ``CodeMode`` capability, or raise if not 
installed.
-
-    Kept here (not a module-level import) because ``pydantic-ai-harness`` is an
-    optional dependency behind the ``code-mode`` extra; importing it eagerly
-    would break installs that don't enable the extra.
-    """
-    try:
-        from pydantic_ai_harness import CodeMode
-    except ImportError as e:
-        # Only report "extra not installed" when pydantic-ai-harness itself is
-        # missing. A failure deeper in its import chain (a broken or missing
-        # transitive dependency) is a different problem -- re-raise it as-is so
-        # the real error isn't masked by a misleading "install the extra" 
message.
-        missing = e.name or ""
-        if missing == "pydantic_ai_harness" or 
missing.startswith("pydantic_ai_harness."):
-            raise AirflowOptionalProviderFeatureException(
-                "code_mode=True requires the 'code-mode' extra. Install it 
with "
-                '`pip install 
"apache-airflow-providers-common-ai[code-mode]"`.'
-            ) from e
-        raise
-    return CodeMode()
-
-
 # CancellableAgentRunMixin must precede BaseOperator so its on_kill overrides 
BaseOperator's
 # no-op. The other mixins only add methods, so they can trail BaseOperator. 
See the MRO guard
 # test in tests/unit/common/ai/mixins/test_cancellable_run.py.
@@ -355,20 +330,8 @@ class AgentOperator(CancellableAgentRunMixin, 
BaseOperator, HITLReviewMixin):
         Cannot be combined with a ``SandboxToolset`` (raises), attached or
         not: a replayed tool result describes a workspace state the replay did
         not reproduce, and the first call that misses the cache runs against
-        whatever the sandbox holds now.
-    :param code_mode: Experimental. When ``True``, wraps the agent's tools in 
a single
-        ``run_code`` tool powered by the Monty sandbox (pydantic-ai-harness
-        ``CodeMode``). Instead of one model round-trip per tool call, the model
-        writes Python that calls the tools as functions, with loops and
-        ``asyncio.gather``, in one turn. The generated code runs in Monty's
-        deny-by-default sandbox; the tools it calls still run in the worker, so
-        ``code_mode`` does not widen what the tools can reach -- it only 
changes
-        how the model invokes them. Requires the ``code-mode`` extra
-        (``pip install "apache-airflow-providers-common-ai[code-mode]"``).
-        Cannot be combined with ``durable=True`` (durable replay assumes a
-        stable per-step call order that code mode does not guarantee), whether
-        code mode comes from this flag or from a ``CodeMode`` capability.
-        Default ``False``.
+        whatever the sandbox holds now. Cannot be combined with a 
pydantic-ai-harness
+        ``CodeMode`` capability (raises).
     :param cache_prompt: When ``True`` (default), asks the provider to cache 
the
         tool definitions, system prompt and conversation so far, so the next
         request in the run -- and a mapped task's other instances within the
@@ -435,10 +398,11 @@ class AgentOperator(CancellableAgentRunMixin, 
BaseOperator, HITLReviewMixin):
     (with the reviewer's reason, when given) and carries on without it. A task
     instance asks at most once per Dag run, across retries and clears; a second
     request fails the task. ``usage_limits`` applies to both sides of the 
pause.
-    Not available together with ``durable``, ``enable_hitl_review``, code mode
-    (``code_mode=True`` or a ``CodeMode`` capability), or a ``SandboxToolset``
+    Not available together with ``durable``, ``enable_hitl_review``, a 
``CodeMode``
+    capability, or a ``SandboxToolset``
     that provisions its own sandbox; there, a tool that requires approval fails
-    the task as before. A ``SandboxToolset`` attached to a
+    the task as before, except one called from inside ``CodeMode``'s 
``run_code``,
+    which does not run and is reported back to the model. A ``SandboxToolset`` 
attached to a
     sandbox another task owns is fine: the sandbox outlives the pause.
 
     :param tool_approval_timeout: Experimental. How long the pause waits for a 
decision.
@@ -493,7 +457,6 @@ class AgentOperator(CancellableAgentRunMixin, BaseOperator, 
HITLReviewMixin):
         agent_params: dict[str, Any] | None = None,
         usage_limits: UsageLimits | dict[str, Any] | None = None,
         durable: bool = False,
-        code_mode: bool = False,
         cache_prompt: bool = True,
         message_history: list[ModelMessage] | str | bytes | None = None,
         # Agent feedback parameters
@@ -528,7 +491,6 @@ class AgentOperator(CancellableAgentRunMixin, BaseOperator, 
HITLReviewMixin):
         self.message_history = message_history
 
         self.durable = durable
-        self.code_mode = code_mode
         self.cache_prompt = cache_prompt
 
         # Populated per run in ``execute`` when durable=True. Declared here so
@@ -556,22 +518,14 @@ class AgentOperator(CancellableAgentRunMixin, 
BaseOperator, HITLReviewMixin):
         if durable and enable_hitl_review:
             raise ValueError("durable=True and enable_hitl_review=True cannot 
be used together.")
 
-        if durable and code_mode:
+        if durable and _contains_code_mode(self._declared_capabilities):
             # Durable replay caches individual model/tool steps via 
CachingModel /
             # CachingToolset and a shared step counter that assumes a stable 
call
             # order across runs. Code mode collapses tools into one 
``run_code``
             # tool and lets the model emit arbitrary Python, so step counts and
             # ordering can differ between the original run and a retry, 
breaking
             # replay. Reject the combination rather than silently 
mis-replaying.
-            raise ValueError("durable=True and code_mode=True cannot be used 
together.")
-
-        if (durable or code_mode) and 
_contains_code_mode(self._declared_capabilities):
-            if durable:
-                # The same conflict as code_mode=True, reached through the 
capability itself.
-                raise ValueError("durable=True cannot be used with a CodeMode 
capability.")
-            # code_mode=True adds a second CodeMode, and pydantic-ai then 
fails the run on a
-            # duplicate ``run_code`` tool without saying where the second one 
came from.
-            raise ValueError("code_mode=True adds a CodeMode capability; pass 
one or the other, not both.")
+            raise ValueError("durable=True cannot be used with a CodeMode 
capability.")
 
         if message_history is not None and enable_hitl_review:
             # The post-review transcript is not recoverable today 
(run_hitl_review
@@ -768,8 +722,6 @@ class AgentOperator(CancellableAgentRunMixin, BaseOperator, 
HITLReviewMixin):
             # ``toolsets=`` wrapping above, so their results would re-execute 
on
             # every retry instead of replaying; wrap their inner toolset too.
             capabilities = self._build_durable_capabilities(capabilities, 
storage, counter)
-        if self.code_mode:
-            capabilities.append(_build_code_mode())
         if self.cache_prompt:
             capabilities.append(PromptCaching())
         if capabilities:
@@ -794,7 +746,6 @@ class AgentOperator(CancellableAgentRunMixin, BaseOperator, 
HITLReviewMixin):
             not AIRFLOW_V_3_3_PLUS
             or self.durable
             or self.enable_hitl_review
-            or self.code_mode
             or _contains_code_mode(self._declared_capabilities)
         ):
             return False
@@ -1229,7 +1180,7 @@ class AgentOperator(CancellableAgentRunMixin, 
BaseOperator, HITLReviewMixin):
             raise UnsupportedToolDeferralError(
                 f"The agent called tools that need approval ({pending_names}), 
but tool approval "
                 "needs Airflow 3.3+ and is not available with durable, 
enable_hitl_review, "
-                "code mode (code_mode=True or a CodeMode capability) or a 
SandboxToolset."
+                "a CodeMode capability or a SandboxToolset."
             )
         store = context["task_state_store"]
         if store.get(_TOOL_APPROVAL_REQUESTED_KEY):
diff --git a/providers/common/ai/tests/unit/common/ai/operators/test_agent.py 
b/providers/common/ai/tests/unit/common/ai/operators/test_agent.py
index 3823eacee0e..01273b99d8a 100644
--- a/providers/common/ai/tests/unit/common/ai/operators/test_agent.py
+++ b/providers/common/ai/tests/unit/common/ai/operators/test_agent.py
@@ -64,7 +64,7 @@ from airflow.providers.common.ai.durable.base import (
 from airflow.providers.common.ai.durable.caching_toolset import CachingToolset
 from airflow.providers.common.ai.durable.step_counter import DurableStepCounter
 from airflow.providers.common.ai.durable.storage import DurableStorage
-from airflow.providers.common.ai.operators.agent import AgentOperator, 
HITLReviewLink, _build_code_mode
+from airflow.providers.common.ai.operators.agent import AgentOperator, 
HITLReviewLink
 from airflow.providers.common.ai.sandbox.base import (
     HOLDER_TAG,
     OWNER_TAG,
@@ -860,8 +860,8 @@ class TestAgentOperatorExecute:
         assert create_call[1]["model_settings"] == {"temperature": 0}
 
     @patch("airflow.providers.common.ai.operators.agent.PydanticAIHook", 
autospec=True)
-    def test_code_mode_default_off_no_capabilities(self, mock_hook_cls, 
make_mock_run_result):
-        """code_mode defaults to False, so with cache_prompt off no 
capabilities are injected."""
+    def test_no_capabilities_injected_with_cache_prompt_off(self, 
mock_hook_cls, make_mock_run_result):
+        """With cache_prompt off and no capabilities passed, create_agent gets 
no capabilities."""
         mock_hook_cls.get_hook.return_value.create_agent.return_value = 
_make_mock_agent(
             "ok", make_mock_run_result
         )
@@ -878,83 +878,6 @@ class TestAgentOperatorExecute:
         create_call = 
mock_hook_cls.get_hook.return_value.create_agent.call_args
         assert "capabilities" not in create_call[1]
 
-    @patch("airflow.providers.common.ai.operators.agent._build_code_mode", 
return_value="CM")
-    @patch("airflow.providers.common.ai.operators.agent.PydanticAIHook", 
autospec=True)
-    def test_code_mode_injects_capability(self, mock_hook_cls, mock_build, 
make_mock_run_result):
-        """code_mode=True appends a CodeMode capability passed to 
create_agent."""
-        mock_hook_cls.get_hook.return_value.create_agent.return_value = 
_make_mock_agent(
-            "ok", make_mock_run_result
-        )
-
-        op = AgentOperator(
-            task_id="t",
-            prompt="hi",
-            llm_conn_id="my_llm",
-            toolsets=[MagicMock(spec=AbstractToolset)],
-            code_mode=True,
-            cache_prompt=False,
-        )
-        op.execute(context=_make_context())
-
-        create_call = 
mock_hook_cls.get_hook.return_value.create_agent.call_args
-        assert create_call[1]["capabilities"] == ["CM"]
-        mock_build.assert_called_once()
-
-    @patch("airflow.providers.common.ai.operators.agent._build_code_mode", 
return_value="CM")
-    @patch("airflow.providers.common.ai.operators.agent.PydanticAIHook", 
autospec=True)
-    def test_code_mode_appends_to_existing_capabilities(
-        self, mock_hook_cls, mock_build, make_mock_run_result
-    ):
-        """A user-supplied capability via agent_params is preserved alongside 
CodeMode."""
-        mock_hook_cls.get_hook.return_value.create_agent.return_value = 
_make_mock_agent(
-            "ok", make_mock_run_result
-        )
-
-        op = AgentOperator(
-            task_id="t",
-            prompt="hi",
-            llm_conn_id="my_llm",
-            code_mode=True,
-            cache_prompt=False,
-            agent_params={"capabilities": ["existing"]},
-        )
-        op.execute(context=_make_context())
-
-        create_call = 
mock_hook_cls.get_hook.return_value.create_agent.call_args
-        assert create_call[1]["capabilities"] == ["existing", "CM"]
-
-    def test_build_code_mode_missing_harness_raises(self):
-        """_build_code_mode raises the optional-feature error when harness is 
absent."""
-        with patch.dict(sys.modules, {"pydantic_ai_harness": None}):
-            with pytest.raises(AirflowOptionalProviderFeatureException, 
match="code-mode"):
-                _build_code_mode()
-
-    def test_build_code_mode_reraises_unrelated_import_error(self):
-        """A broken transitive import inside the harness is re-raised, not 
masked as 'extra missing'."""
-        real_import = __import__
-
-        def fake_import(name, *args, **kwargs):
-            if name == "pydantic_ai_harness":
-                raise ModuleNotFoundError("No module named 'a_broken_dep'", 
name="a_broken_dep")
-            return real_import(name, *args, **kwargs)
-
-        with patch("builtins.__import__", side_effect=fake_import):
-            with pytest.raises(ModuleNotFoundError, match="a_broken_dep"):
-                _build_code_mode()
-
-    @patch("airflow.providers.common.ai.operators.agent._build_code_mode")
-    def test_code_mode_not_built_at_init(self, mock_build):
-        """code_mode is serialization-safe: the CodeMode capability is built 
lazily in
-        _build_agent, never at construction time (so nothing non-serializable 
is stored)."""
-        op = AgentOperator(task_id="t", prompt="hi", llm_conn_id="my_llm", 
code_mode=True)
-        mock_build.assert_not_called()
-        assert op.code_mode is True
-
-    def test_durable_and_code_mode_rejected(self):
-        """durable and code_mode cannot be combined (durable replay assumes 
stable step order)."""
-        with pytest.raises(ValueError, match="durable=True and 
code_mode=True"):
-            AgentOperator(task_id="t", prompt="hi", llm_conn_id="my_llm", 
durable=True, code_mode=True)
-
     @patch("airflow.providers.common.ai.operators.agent.PydanticAIHook", 
autospec=True)
     def test_cache_prompt_default_on_appends_capability_last(self, 
mock_hook_cls, make_mock_run_result):
         """cache_prompt defaults to True and adds PromptCaching after any user 
capability."""
@@ -1366,12 +1289,6 @@ class TestAgentOperatorCapabilities:
 
         assert op._supports_tool_approval() is True
 
-    def 
test_code_mode_flag_and_code_mode_capability_are_refused_together(self, 
fake_harness):
-        with pytest.raises(ValueError, match="one or the other"):
-            AgentOperator(
-                task_id="t", prompt="p", llm_conn_id="llm", code_mode=True, 
capabilities=[_FakeCodeMode()]
-            )
-
     def test_capability_function_is_left_for_the_run_to_resolve(self, 
fake_harness):
         def build(ctx):
             return Thinking()
diff --git 
a/providers/common/ai/tests/unit/common/ai/operators/test_agent_tool_approval.py
 
b/providers/common/ai/tests/unit/common/ai/operators/test_agent_tool_approval.py
index 09e2fe93732..7334b98fc2e 100644
--- 
a/providers/common/ai/tests/unit/common/ai/operators/test_agent_tool_approval.py
+++ 
b/providers/common/ai/tests/unit/common/ai/operators/test_agent_tool_approval.py
@@ -551,7 +551,6 @@ class TestWhenApprovalApplies:
         "kwargs",
         [
             pytest.param({"durable": True}, id="durable"),
-            pytest.param({"code_mode": True}, id="code_mode"),
             pytest.param({"enable_hitl_review": True}, id="hitl_review"),
             pytest.param({"toolsets": [SandboxToolset(_NoopBackend())]}, 
id="sandbox"),
             pytest.param(
diff --git a/uv.lock b/uv.lock
index f4bc231f9d3..ce3eafa7ab1 100644
--- a/uv.lock
+++ b/uv.lock
@@ -4763,7 +4763,7 @@ requires-dist = [
     { name = "opensandbox", marker = "extra == 'opensandbox'", specifier = 
">=1.1.0" },
     { name = "pyarrow", marker = "python_full_version >= '3.14' and extra == 
'parquet'", specifier = ">=22.0.0" },
     { name = "pyarrow", marker = "python_full_version < '3.14' and extra == 
'parquet'", specifier = ">=18.0.0" },
-    { name = "pydantic-ai-harness", extras = ["codemode"], marker = "extra == 
'code-mode'", specifier = ">=0.3.0" },
+    { name = "pydantic-ai-harness", extras = ["codemode"], marker = "extra == 
'code-mode'", specifier = ">=0.24.0" },
     { name = "pydantic-ai-shields", marker = "extra == 'shields'", specifier = 
">=0.3.4" },
     { name = "pydantic-ai-skills", marker = "extra == 'skills'", specifier = 
">=1.2.0" },
     { name = "pydantic-ai-slim", specifier = ">=2.33.0" },

Reply via email to