This is an automated email from the ASF dual-hosted git repository.
kaxil pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/airflow.git
The following commit(s) were added to refs/heads/main by this push:
new 0609550c022 Lead the `common.ai` sandbox docs with when to use it and
where each piece runs (#74297)
0609550c022 is described below
commit 0609550c0221a087109abece8f5c1e2c7bdd6290
Author: Kaxil Naik <[email protected]>
AuthorDate: Tue Oct 6 13:52:17 2026 +0100
Lead the `common.ai` sandbox docs with when to use it and where each piece
runs (#74297)
The sandbox page opened on the problem and only answered "should I use this,
or the executor, the pod operator or code mode?" across four sections
further
down. It now opens with that answer, draws what moves into the sandbox and
what stays on the worker, merges the two overlapping limitation lists, and
shows the task log and XCom of a real run of the quick-start agent on Modal.
---
providers/common/ai/docs/concepts.rst | 2 +-
providers/common/ai/docs/sandbox/backends.rst | 2 +
providers/common/ai/docs/sandbox/index.rst | 416 +++++++++++++-------------
3 files changed, 208 insertions(+), 212 deletions(-)
diff --git a/providers/common/ai/docs/concepts.rst
b/providers/common/ai/docs/concepts.rst
index af5d9c0c26a..27c64b05426 100644
--- a/providers/common/ai/docs/concepts.rst
+++ b/providers/common/ai/docs/concepts.rst
@@ -49,7 +49,7 @@ generating SQL, comparing schemas, or submitting many prompts
as one batch.
The AI step is orchestrated by Airflow: the model calls, the agent loop, and
any tools
run in the Airflow worker by default, where they get retries, logging, and
observability like
-any other task. The exception is :ref:`SandboxToolset <sandbox-limitations>`,
which exists so
+any other task. The exception is :ref:`SandboxToolset <sandbox-placement>`,
which exists so
that code the *model* writes runs somewhere else.
Toolsets give agents reach
diff --git a/providers/common/ai/docs/sandbox/backends.rst
b/providers/common/ai/docs/sandbox/backends.rst
index acf3d453938..9fa9369e27e 100644
--- a/providers/common/ai/docs/sandbox/backends.rst
+++ b/providers/common/ai/docs/sandbox/backends.rst
@@ -248,6 +248,8 @@ The runtime remains a deployment choice. The default Docker
runtime shares the
host kernel; choose a stronger runtime such as Kata when your threat model
needs
a VM boundary.
+.. _sandbox-backend-sbx:
+
sbx (Docker Sandboxes, local)
-----------------------------
diff --git a/providers/common/ai/docs/sandbox/index.rst
b/providers/common/ai/docs/sandbox/index.rst
index 38fce059808..b32b9abc7eb 100644
--- a/providers/common/ai/docs/sandbox/index.rst
+++ b/providers/common/ai/docs/sandbox/index.rst
@@ -25,15 +25,78 @@ Sandboxed execution for agents
Experimental: this can change or be removed in a minor release of this
provider.
See :ref:`howto/stability`.
-An agent that is asked to do open-ended work writes code, and then something
has
-to run that code. By default that something is the Airflow worker: a skill
-script, a hand-written tool that shells out, or generated glue in
-:ref:`code mode <code-mode>` all execute on the host that also holds your
-connections, your filesystem and your network position, and the code they run
-was generated a moment ago from inputs you do not control.
+:class:`~airflow.providers.common.ai.toolsets.sandbox.SandboxToolset` gives an
agent a
+disposable workspace to run the code it writes. Without one, that code runs on
the
+Airflow worker: a skill script, a hand-written tool that shells out, or
generated glue
+in :ref:`code mode <code-mode>` all execute on the host that holds your
connections,
+your filesystem and your network position, and the code was generated a moment
ago
+from inputs you do not control.
+
+Use it when a model inside a running task writes code you cannot name in
advance, and
+that code needs a real shell, a real interpreter or packages. Use something
else when:
+
+- **The model only calls operations you can name.** Use
+ :class:`~airflow.providers.common.ai.toolsets.sql.SQLToolset` or
+ :class:`~airflow.providers.common.ai.toolsets.hook.HookToolset`. They are
narrower,
+ their results are bounded, and the credential never reaches the model.
+- **The model chains tools you registered with a little glue logic**, such as
looking
+ up 200 customers and flagging the ones with an open ticket. Use
+ :ref:`code mode <code-mode>`. It needs no backend and starts in well under a
+ millisecond, against roughly a second to provision a hosted sandbox once its
image
+ is built.
+- **The whole task, connections included, has to run off a shared worker.** Use
+ ``KubernetesExecutor``, with a pod override for one team's tasks. It
isolates the
+ task from your platform, not the task from the code its model writes, so the
two
+ combine (:ref:`sandbox-boundaries`).
+- **You are launching a payload whose image and command are known when it
starts**,
+ such as a vendor's scanner image, or even a complete agent. Use
+ ``KubernetesPodOperator``.
+
+.. _sandbox-placement:
+
+Where each piece runs
+---------------------
-:class:`~airflow.providers.common.ai.toolsets.sandbox.SandboxToolset` gives the
-model a disposable workspace for that code instead. It exposes four tools:
+The toolset moves exactly four tools into the sandbox. Everything else on the
agent
+stays where the task runs:
+
+.. mermaid::
+
+ flowchart TB
+ model["Model provider"]
+ db[("Warehouse")]
+ subgraph host["Worker host or task pod"]
+ subgraph task["Airflow task"]
+ loop["Agent loop<br/>message history, LLM credential"]
+ other["SQLToolset, HookToolset and every other
toolset<br/>connection credentials stay here"]
+ client["SandboxToolset<br/>backend credential stays here"]
+ end
+ sbx["sbx microVM<br/>local development only"]
+ end
+ remote["Modal or OpenSandbox sandbox<br/>run_command,
read_file,<br/>write_file, list_directory"]
+ model <--> loop
+ db <--> other
+ loop <--> other
+ loop <--> client
+ client <--> remote
+ client <-.->|"or, on a laptop"| sbx
+
+Commands and files the model sends through these four tools run inside the
sandbox,
+with only what ``SandboxSpec.env`` puts there. No connection, variable or
worker
+environment variable goes in, and neither does the credential that provisioned
the
+sandbox. Tool results come back to the agent loop and from there to the model.
The
+agent loop, the model calls and every other toolset keep the task's authority,
so the
+sandbox contains model-written code, not the agent (:ref:`sandbox-security`).
Under
+``KubernetesExecutor`` the outer box is the task's pod: the pod protects the
cluster
+from the task, and the sandbox protects the task from its model. An ``sbx``
microVM
+cannot run inside an unprivileged pod, so that combination needs Modal or
OpenSandbox.
+
+The credential that provisions the sandbox depends on the backend: a ``modal``
+connection (``modal_conn_id``; see the :ref:`Modal connection page
<howto/connection:modal>`)
+for Modal; ``opensandbox_conn_id``, or ``OPEN_SANDBOX_DOMAIN`` and
``OPEN_SANDBOX_API_KEY``,
+for OpenSandbox; and the host's ``sbx login`` for ``sbx``.
+
+The four tools:
.. list-table::
:widths: 25 75
@@ -53,22 +116,20 @@ model a disposable workspace for that code instead. It
exposes four tools:
* - ``list_directory``
- Lists a directory. Directories are shown with a trailing ``/``.
-The sandbox is provisioned by a
-:class:`~airflow.providers.common.ai.sandbox.SandboxBackend` on the model's
first
-tool call and torn down when the agent run ends. Three backends ship: a hosted
one on
-`Modal <https://modal.com/docs/guide/sandbox>`__ and a self-hosted one on
-`OpenSandbox <https://open-sandbox.ai/>`__ for production and Kubernetes, and
a local
-microVM one on `Docker Sandboxes <https://docs.docker.com/ai/sandboxes/>`__ for
-development. The four tool names and shapes match pydantic-ai's own sandbox
-capabilities, so a model that has seen one already knows this one.
-
-**Adding this toolset gives the agent shell and file operations in a separate
-workspace. Every other tool keeps its existing permissions and runs where it
ran
-before.** Whether that is a new capability depends on what the agent already
had:
-an agent with a skill script or a shell tool already ran generated code on the
-worker, and this is where that code should run instead. An agent that only had
-named database operations is being granted a general shell for the first time,
in
-a workspace with less authority than the worker, and you are choosing to grant
it.
+A :class:`~airflow.providers.common.ai.sandbox.SandboxBackend` provisions the
sandbox
+on the model's first tool call and tears it down when the agent run ends. Three
+backends ship: `Modal <https://modal.com/docs/guide/sandbox>`__ (hosted) and
+`OpenSandbox <https://open-sandbox.ai/>`__ (self-hosted) for production and
+Kubernetes, and `Docker Sandboxes <https://docs.docker.com/ai/sandboxes/>`__
(``sbx``,
+a microVM on the worker host) for local development. The tool names are the
ones
+pydantic-ai's `Shell <https://pydantic.dev/docs/ai/harness/shell/>`__ and
+`FileSystem <https://pydantic.dev/docs/ai/harness/filesystem/>`__ capabilities
use.
+
+**Whether this grants something new depends on what the agent already had.**
An agent
+with a skill script or a shell tool already ran generated code on the worker,
and this
+is where that code should run instead. An agent that only had named database
+operations is being granted a general shell for the first time, in a workspace
with
+less authority than the worker, and you are choosing to grant it.
Pages in this section
---------------------
@@ -96,6 +157,35 @@ and a bad pivot or a runaway loop costs a small sandbox
rather than a worker slo
:start-after: [START howto_sandbox_agent_investigation]
:end-before: [END howto_sandbox_agent_investigation]
+Run against a test warehouse with two weeks of daily revenue, in which EMEA
web revenue
+was seeded to jump on the run date, the task log ends with a summary of the
run and the
+order of the agent's tool calls. Every ``run_command`` and ``write_file`` ran
in the Modal
+sandbox; every ``query`` ran in the task:
+
+.. code-block:: text
+
+ LLM run complete: model=claude-sonnet-5, requests=27, tool_calls=27,
input_tokens=273753, output_tokens=7137, total_tokens=280890
+ LLM run cost: $0.1557896 (USD, best-effort)
+ Tool call sequence: list_tables -> get_schema -> get_schema -> query ->
write_file -> write_file -> run_command -> run_command -> query -> query ->
query -> query -> write_file -> run_command -> query -> query -> query -> query
-> run_command -> run_command -> run_command -> run_command -> run_command ->
run_command -> run_command -> write_file -> run_command -> final_result
+
+The ``Findings`` the task returned to XCom, with ``summary`` and
``suspected_cause``
+shortened here:
+
+.. code-block:: json
+
+ {
+ "summary": "Total revenue on 2026-10-05 was $24,864 vs a trailing-7-day
average (09-28 to 10-04) of $21,522 — a +15.5% move, exceeding the 10%
threshold. Breaking revenue down by region×channel, EMEA/web is the sole driver
...",
+ "confidence": "high",
+ "suspected_cause": "A step-change in average order value for EMEA/web
orders on 2026-10-05 (order count stayed flat at 3, but per-order amount jumped
from ~$1,600-$1,733 to $3,040, ~1.75x normal) ...",
+ "affected_segments": ["EMEA / web"]
+ }
+
+An open-ended investigation can take many model requests. With
``usage_limits`` unset,
+pydantic-ai's own default of 50 requests still applies, and a run that needs
more fails
+the task with ``UsageLimitExceeded: The next request would exceed the
request_limit of
+50``; one of the two runs of this Dag captured for this page did. Pass
``usage_limits={"request_limit": 100}``, or
+a ``cost_limit``, to set the budget yourself.
+
The same toolset on the local ``sbx`` backend, for developing on a laptop. Here
the agent maps a vendor file whose columns drift onto a staging schema: it
writes a
script, runs it, reads the traceback and fixes it, which is the loop nobody can
@@ -107,6 +197,15 @@ is the only change between this Dag and production:
:start-after: [START howto_sandbox_agent_local]
:end-before: [END howto_sandbox_agent_local]
+``host_network_policy="deny-all"`` is there because ``sbx`` cannot block
egress per
+sandbox. Leave it out and the backend refuses to provision the default spec,
which
+denies all egress, and the task fails
+(:ref:`a failed provisioning fails the task <sandbox-lifecycle>`) with:
+
+.. code-block:: text
+
+ airflow.providers.common.ai.sandbox.base.SandboxTerminalError: SandboxSpec
asks for no network egress, but this backend cannot enforce that per sandbox
and the host policy has not been declared. Run 'sbx policy init deny-all' on
the worker host and pass host_network_policy='deny-all', or pass
SandboxSpec(block_network=False) to acknowledge that egress is open.
+
Install the Modal extra, which also installs the Modal provider, and create a
``modal`` connection with the Modal token id as its login and the token secret
as
its password. The backend reads ``modal_default`` unless you pass
``modal_conn_id``;
@@ -124,10 +223,8 @@ credentials present:
What it is for, in practice
---------------------------
-The rule is: a model inside a running task writes code you cannot name in
advance,
-and that code needs a real shell, a real interpreter or packages. These are the
-jobs that rule describes, each paired with the version where a sandbox is the
wrong
-call.
+Four jobs where the model writes code nobody can write down in advance, each
paired
+with the version where a sandbox is the wrong call.
**A vendor drops a CSV nobody has a schema for.** Three date formats, an
encoding
error at row 41,000, pandas needed. Every step depends on the previous step's
@@ -146,8 +243,8 @@ carrying them through the model's context.
**Migrate forty dbt models between warehouse dialects.** The model files go
into
the workspace, the image carries ``sqlglot`` and ``sqlfluff``, and the loop of
convert, lint, read errors, fix runs for as long as it takes, fully offline.
The
-rewritten files are the deliverable, so the sandbox is one a task provisions
and
-reads out afterwards (:ref:`sandbox-attach`).
+rewritten files are the deliverable, so they leave through ``exports`` when
the run
+ends (:ref:`sandbox-results`).
**A failing task's traceback, handed to an agent to propose a fix.** The agent
reproduces the crash against a sample, tries a fix, reruns, and posts a diff
for a
@@ -155,28 +252,14 @@ person to review. That code ran with the worker's
credentials the first time; th
reproduction should not. A written diagnosis with no execution is
``LLMOperator`` reading the log.
-**Not a sandbox: flag 200 customers with an open ticket** using two existing
-tools. That is :ref:`code mode <code-mode>`: a short loop over registered
tools,
-no packages, no shell, sub-millisecond start.
-
-**Not a sandbox: run a vendor's PII-scanner image over an S3 prefix.** No model
-writes anything. That is ``KubernetesPodOperator``, a fixed payload with
templated
-arguments in its own pod.
-
-**Not only a sandbox: one team's agent tasks must run off the shared workers.**
-That is a where-does-the-task-run requirement, so ``KubernetesExecutor`` with a
-pod override, and the sandbox toolset inside the pod if those agents also write
-code. The pod protects the cluster from the task; the sandbox protects the task
-from its model.
-
.. _sandbox-boundaries:
Choosing the boundary
---------------------
"Sandboxed" means different things at different layers, and picking the wrong
-layer is the most common way to end up with less protection than you think.
Four
-boundaries exist, from smallest to largest:
+layer leaves you with less protection than you think. Four boundaries exist,
from
+smallest to largest:
.. list-table::
:widths: 22 33 45
@@ -206,40 +289,14 @@ boundaries exist, from smallest to largest:
- Everything outside the task. Not the task from its own code: the
agent and what it runs are both inside this boundary.
-``SandboxToolset`` is the first row. It relocates exactly its four tools and
-nothing else. Boundary size is not a security level on its own: the image, the
-credentials you inject, the network policy and the resource limits decide the
-actual isolation. Choose the smallest boundary that fits, then configure it.
-
-.. list-table::
- :widths: 48 52
- :header-rows: 1
-
- * - The agent needs to
- - Reach for
- * - Call operations you can name in advance
- - :class:`~airflow.providers.common.ai.toolsets.sql.SQLToolset` or
- :class:`~airflow.providers.common.ai.toolsets.hook.HookToolset`. Narrow,
- bounded, and the credential stays in the worker rather than reaching the
- model. If you can name the operations, name them.
- * - Chain those tools with glue logic
- - :ref:`code mode <code-mode>`. Monty confines generated code to the tools
- you registered and starts in well under a millisecond, against roughly a
- second to provision a hosted sandbox. Reach past it for a sandbox when
the
- model needs something Monty does not have: a shell, ``pip install``, a
- compiled library.
- * - Write and run open-ended code: reshape data with no known schema,
install
- a package, or fix its own failing script by reading the traceback
- - ``SandboxToolset``
- * - Produce a large artifact for a downstream task
- - A ``@task`` driving a backend directly when the Dag knows the job. When
- the agent has to produce it, ``SandboxToolset(exports=...)`` copies the
file
- to object storage when the run ends; see :ref:`sandbox-results`.
- * - A whole task's worth of untrusted work isolated, with no agent involved
- - ``KubernetesPodOperator``
- * - Airflow's own credentials kept away from the agent
- - Not this. Scope the connections the task can see, and give the agent
only
- the toolsets it needs. See :ref:`sandbox-security`.
+``SandboxToolset`` is the first row. Boundary size is not a security level on
its
+own: the image, the credentials you inject, the network policy and the resource
+limits decide the actual isolation. Choose the smallest boundary that fits,
then
+configure it. Two needs are not a boundary at all. A large file for a
downstream task
+leaves through ``exports`` or through a task that owns the sandbox
+(:ref:`sandbox-results`). Keeping Airflow's own credentials away from the
agent is done
+by scoping the connections the task can see and the toolsets the agent gets
+(:ref:`sandbox-security`).
**Code mode and the sandbox compose.** With both enabled, the three file tools
fold into ``run_code`` as callables the generated code can loop over, and
@@ -252,9 +309,8 @@ reach the same sandbox: there is one per agent run
whichever way a call arrives.
What enforces the boundary
^^^^^^^^^^^^^^^^^^^^^^^^^^
-Swapping ``SbxSandboxBackend`` for ``ModalSandboxBackend`` changes how the
tool-call
-boundary is enforced and leaves its contents alone: the same four tools run in
the
-sandbox, and every other toolset stays on the worker.
+Swapping one backend for another changes how the tool-call boundary is
enforced and
+leaves its contents alone: the same four tools run in the sandbox.
.. list-table::
:widths: 20 28 26 26
@@ -282,7 +338,8 @@ sandbox, and every other toolset stays on the worker.
- Left running. No server-side lifetime, so the microVM and its workspace
directory survive.
* - ``ModalSandboxBackend``
- - Modal's container runtime, which Modal documents as gVisor.
+ - Modal's container runtime, which Modal
+ `documents as gVisor <https://modal.com/docs/guide/security>`__.
- Modal's infrastructure, off the worker.
- Ended by Modal at ``sandbox_timeout``, or ``idle_timeout`` if set.
* - ``OpenSandboxBackend``
@@ -294,9 +351,9 @@ sandbox, and every other toolset stays on the worker.
When a run ends normally, the task calls the backend's ``destroy``. ``sbx``
runs its
removal command and waits up to two minutes for it; Modal and OpenSandbox each
send a
termination request and return without waiting for the sandbox to stop. Any of
them
-can return with the sandbox still present, and none of those cases fails the
task. A SIGKILL, an
-out-of-memory kill or a lost node skips that teardown entirely, and then only
the
-last column applies.
+can return with the sandbox still present, and none of those cases fails the
task. A
+SIGKILL, an out-of-memory kill or a lost node skips that teardown entirely,
and then
+only the last column applies.
Why not ``KubernetesPodOperator`` or the ``KubernetesExecutor``?
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
@@ -308,9 +365,7 @@ writes.
Run an ``AgentOperator`` under the ``KubernetesExecutor`` and the whole task,
agent loop included, gets its own pod. Model-written code still executes inside
that pod, with its service account, its mounted secrets, its connections and
its
-network position. The pod protects the cluster from the task; it does not
protect
-the task from what the model decided to run, because from the pod's point of
view
-that code is the task.
+network position. From the pod's point of view that code is the task.
``KubernetesPodOperator`` does contain the work, and its image, command,
arguments
and environment are all templated, so it can launch a different payload per
run,
@@ -343,102 +398,70 @@ whatever it wants inside the pod that also holds your
credentials.
- Only what ``SandboxSpec.env`` names, which is nothing by default
Running an agent in a pod *and* giving it a sandbox is the setup a Kubernetes
-deployment usually wants: the pod bounds the task, the sandbox bounds the
model.
-The ``sbx`` backend cannot be the second half of that, because it needs KVM on
the
-worker host and an unprivileged pod cannot provide it; that is what the hosted
-backend is for.
+deployment usually wants, with Modal or OpenSandbox as the backend.
When to choose it
-----------------
**Choose it when** the work is running code the model wrote, rather than
calling
a tool you picked in advance: exploratory analysis, installing a package for
one
-task, a script the model writes, runs, and fixes from its own traceback. Every
-other toolset route answers "call this thing"; this one answers "here is
-somewhere to work". This page has the worked scenarios above and
-the limitations below to read before designing a Dag around it.
-
-Before reaching for it, check whether the actual need is narrower than that.
-:ref:`Code mode <code-mode>` is a capability on ``AgentOperator``. It changes
how the
-model invokes the tools it already has, letting it write code that calls
several
-of them instead of emitting one call per step. It does not give the agent
somewhere
-to run arbitrary code of its own. Because it needs no backend, it avoids the
-production-readiness, network-isolation and reclamation caveats of the ``sbx``
-backend on :doc:`backends`. It does not change what the agent can reach: the
-generated code runs in Monty's deny-by-default sandbox, but the tools it calls
-still run in the worker, so a credential-bearing toolset on the same agent
stays
-within reach whether or not code mode is on. See :ref:`code-mode` and
-:ref:`sandbox-boundaries`.
+task, a script the model writes, runs, and fixes from its own traceback.
+
+.. _sandbox-limitations:
**What it cannot do**
-- One of its three backends does not run on Kubernetes. ``SbxSandboxBackend``
drives
- Docker Sandboxes on the worker host, and its own documentation says to use it
- for local development: it wants the ``sbx`` binary on the host, an
- authenticated Docker account, a one-time ``sbx policy init``, and on Linux
KVM
- or nested virtualization, which an unprivileged container cannot provide.
- Production and Kubernetes use a remote backend instead, either
- :class:`~airflow.providers.common.ai.sandbox.modal.ModalSandboxBackend`
behind
- the ``modal`` extra for a managed service, or
- :class:`~airflow.providers.common.ai.sandbox.opensandbox.OpenSandboxBackend`
- behind ``opensandbox`` for a self-hosted one. Neither installs anything on
the
- worker, and each reclaims a sandbox at its own server-side lifetime if the
- worker dies. All three implement
- :class:`~airflow.providers.common.ai.sandbox.SandboxBackend`, and another
- vendor can too.
-- It does not contain the agent. Only what these tools do runs in the sandbox;
- the agent loop, the model calls, and every other toolset on the same agent
stay
- in the worker with the worker's credentials. It contains model-written code,
so
- pairing it with a credential-bearing toolset on the same agent puts the
- credential back within reach. :ref:`sandbox-boundaries` sets this out in
full.
-- The ``sbx`` backend cannot enforce network isolation on its own. It applies
- ``allow_egress_to`` as a per-sandbox policy rule, but only on top of a
- ``deny-all`` host policy, since a local rule can narrow egress and never
- widen it; ``block_network`` has no per-sandbox enforcement at all. Ask for
- either against a host policy that is not already ``deny-all`` and the
- backend raises rather than silently leaving the sandbox less restricted
- than the spec asked for. ``block_network`` defaults to ``True``, so a bare
- ``SandboxSpec()`` with no arguments already asks for it and is refused
- under the default ``host_network_policy="unknown"``. On Modal the same
default
- maps onto the sandbox's own ``block_network`` and is enforced exactly; a
- hostname allowlist there is matched on the TLS handshake name and has to be
- opted into, for the reasons set out on :doc:`backends`. OpenSandbox enforces
- hostname allowlists with its egress sidecar, but refuses
- ``allow_egress_to_cidrs`` because the SDK cannot prove the sidecar is in the
- ``dns+nft`` mode required for CIDR enforcement.
-- Reclamation depends on the backend. A failed teardown is logged as a warning
- rather than raised, deliberately, so that a teardown blip cannot fail a
- finished run. On ``sbx`` nothing else picks up the slack: there is no
- server-side TTL, so a worker killed outright leaves the microVM and its
- workspace directory behind, named ``airflow-sandbox-*`` so an operator can
find
- and remove them. On Modal the sandbox ends at its own ``sandbox_timeout``
- whatever became of the worker.
-- A sandbox the toolset provisions itself lives for one run, and a file the
- agent built in it can leave only through the model's context, which is
text-only
- and capped. When a file has to come out, or a credential has to come from a
- connection, a ``@task`` provisions the sandbox and the agent attaches to it;
- :ref:`sandbox-attach` has the example. When the Dag already knows the job, a
- ``@task`` drives the backend and no agent is involved.
-
-**A real example.** ``example_sandbox_toolset.py`` in this provider's example
-Dags has an agent investigating a revenue anomaly on the Modal backend beside a
-``SQLToolset``, the same agent shape on ``sbx`` for a laptop, a ``@task``
-producing a file through a sandbox, and a task-owned sandbox that an agent
attaches
-to; all four are on this page and :doc:`configuration`. The two
-system tests, ``example_sandbox_toolset_sbx.py`` and
-``example_sandbox_toolset_modal.py``, run against a real backend and are
+- **It does not contain the agent.** Pairing it with a credential-bearing
toolset
+ on the same agent leaves that credential within reach of a steered model.
+ :ref:`sandbox-security`.
+- **Nothing survives the run** in a sandbox the toolset provisions itself,
+ including across task retries. On Modal, a sandbox a task provisions and the
agent
+ attaches to does; ``sbx`` and OpenSandbox cannot be attached to.
+ :ref:`Lifecycle <sandbox-lifecycle>`,
+ :ref:`A sandbox another task owns <sandbox-attach>`.
+- **A file the agent built leaves through** ``exports`` **or a task**, never
through
+ the model's context, which is text-only and capped.
+ :ref:`Getting a result out <sandbox-results>`.
+- **A credential handed to the code inside the sandbox comes from a connection
+ only when a task provisions the sandbox**, which needs Modal; the toolset's
own spec
+ is fixed at parse time, and anything injected is readable by the model.
:ref:`Credentials <sandbox-credentials>`.
+- **It cannot be combined with** ``durable=True``. ``enable_hitl_review=True``
and
+ per-tool approval work only when the sandbox is task-owned:
``AgentOperator`` refuses
+ HITL review beside a sandbox the toolset provisions itself, and a tool that
requires
+ approval there fails the task. :ref:`Lifecycle <sandbox-lifecycle>`.
+- **A run that outlives** ``sandbox_timeout`` **fails the task.**
+ :ref:`Lifecycle <sandbox-lifecycle>`.
+- **Network rules are only as strong as the backend.** A backend that cannot
enforce
+ a network field raises rather than provisioning a looser sandbox, so a bare
+ ``SandboxSpec()`` is refused on ``sbx`` until the host policy is declared
+ ``deny-all``. Modal's hostname allowlist is a weak control and has to be
opted into;
+ its address allowlist is enforced properly but cannot serve a package
registry whose
+ addresses rotate. :ref:`Configuring a sandbox <sandbox-configuring>`,
+ :ref:`Modal <sandbox-backend-modal>`.
+- **On Modal, commands run as root and** ``workdir`` **is not a jail.**
+ :ref:`Modal <sandbox-backend-modal>`.
+- **The** ``sbx`` **backend is for local development.** It needs the ``sbx``
binary,
+ a Docker login and, on Linux, KVM or nested virtualization, so it does not
run on
+ unprivileged Kubernetes, and a worker killed outright leaves its microVM
behind.
+ :ref:`sbx <sandbox-backend-sbx>`.
+- **A failed teardown is logged, not raised**, so a teardown blip cannot fail a
+ finished run; reclaiming the sandbox is then the backend's lifetime or an
+ operator's sweep. :ref:`Cost and operations <sandbox-cost>`.
+
+**A real example.** ``example_sandbox_toolset.py`` in this provider's example
Dags
+has the two agents in the quick start, a ``@task`` producing a file through a
+sandbox, an agent attaching to a task-owned sandbox, and an agent exporting a
file,
+shown on this page and in :doc:`configuration`. The system tests
+``example_sandbox_toolset_sbx.py``, ``example_sandbox_toolset_modal.py`` and
+``example_sandbox_toolset_opensandbox.py`` run against a real backend and are
reachable from the System Tests entry in the sidebar.
-**Credentials and where it runs.** Airflow puts none of its context,
connections,
-variables or worker environment into the sandbox; only what you pass through
-:class:`~airflow.providers.common.ai.sandbox.SandboxSpec` goes in, and the
-credential that provisions the sandbox never enters it. For Modal that
credential
-is a ``modal`` connection (``modal_conn_id``; see the
-:ref:`Modal connection page <howto/connection:modal>`). For ``sbx`` it sits
outside Airflow:
-``sbx login`` on the machine. Work runs in a per-run microVM on the worker
host with
-``sbx``, or off the worker entirely in Modal's infrastructure. Its tool calls
-act as barriers, as they do for the other
-routes that build their own tools; see :ref:`toolset-call-barriers`.
+**Credentials and where it runs.** Only what ``SandboxSpec.env`` names enters
the
+sandbox; the provisioning credential for each backend is listed under
+:ref:`sandbox-placement`. Work runs in Modal's infrastructure, on the
OpenSandbox
+server's runtime (off the worker unless you run the server there), or in a
per-run
+microVM on the worker host with ``sbx``. Its tool calls act as barriers, as
they do
+for the other routes that build their own tools; see
:ref:`toolset-call-barriers`.
.. _sandbox-other-frameworks:
@@ -469,42 +492,12 @@ block ends, however the agent finishes:
agent = Agent(plugins=[AirflowTools(warehouse, sandbox)])
return str(agent("Reconcile the September ledger against the
warehouse."))
-A tool call made before the block or after it is refused rather than
provisioning a
-sandbox nothing would destroy. The toolset's own error rules hold as well: a
command
+A tool call made before the block or after it raises ``SandboxTerminalError``
saying
+the toolset is not open and must be used inside ``with sandbox:`` or
+``async with sandbox:``, rather than provisioning a sandbox nothing would
destroy. The toolset's own error rules hold as well: a command
that fails is output the model reads, and a sandbox that cannot be provisioned
ends the
agent run so the task fails.
-.. _sandbox-limitations:
-
-Limitations
------------
-
-These apply to every backend. Each is explained in the section it belongs to;
this
-is the list to read before designing a Dag around an agent with a sandbox.
-
-- **Nothing survives the run** in a sandbox the toolset provisions itself,
- including across task retries. A sandbox a task provisions and the agent
- attaches to does. :ref:`Lifecycle <sandbox-lifecycle>`,
- :ref:`A sandbox another task owns <sandbox-attach>`.
-- **A file the agent built leaves only through a task**, never through the
- model's context, which is text-only and capped.
- :ref:`Getting a result out <sandbox-results>`.
-- **A credential handed to the code inside the sandbox comes from a connection
- only when a task provisions the sandbox**; the toolset's own spec is fixed at
- parse time, and anything injected is readable by the model.
:ref:`Credentials <sandbox-credentials>`.
-- **Cannot be combined with** ``durable=True``, and with
``enable_hitl_review=True``
- only when the sandbox is task-owned; ``AgentOperator`` raises otherwise.
- :ref:`Lifecycle <sandbox-lifecycle>`.
-- **A run that outlives** ``sandbox_timeout`` **fails the task.**
- :ref:`Lifecycle <sandbox-lifecycle>`.
-- **The hostname allowlist is a weak control** and refused unless opted into.
The
- address allowlist is enforced properly but cannot serve a package registry
whose
- addresses rotate.
- :ref:`Modal <sandbox-backend-modal>`.
-- **Commands run as root and** ``workdir`` **is not a jail.**
- :ref:`Modal <sandbox-backend-modal>`.
-- **It does not contain the agent.** :ref:`sandbox-security`.
-
.. _sandbox-security:
Security boundary
@@ -523,8 +516,8 @@ The controls for that risk are different controls. Give the
agent only the
toolsets it needs. Point every connection at a least-privilege role. Prefer a
capability toolset, which lets the model *use* a connection without ever
holding
it, over a credential in ``SandboxSpec.env``, which the model can read. And
treat
-a table allowlist as a guardrail that contains intent rather than a boundary
that
-contains access.
+a table allowlist as a check on which tables the model asks for, not a limit
on what
+the connection can reach.
Isolation is one of several controls an agent needs, and each has a limit:
@@ -565,7 +558,8 @@ Isolation is one of several controls an agent needs, and
each has a limit:
- Reviewing the final answer does not undo writes made during the run.
- Keep writes out of the agent and gate the task that makes them.
``AgentOperator`` refuses ``enable_hitl_review`` beside a
- ``SandboxToolset`` it can inspect; see :ref:`Lifecycle
<sandbox-lifecycle>`.
+ ``SandboxToolset`` that provisions its own sandbox; see
+ :ref:`Lifecycle <sandbox-lifecycle>`.
* - Deadlines, resource limits, cleanup
- What bounds runaway work and orphaned resources?
- A resource request or a cost report is not an enforced ceiling.