This is an automated email from the ASF dual-hosted git repository.
kaxil pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/airflow.git
The following commit(s) were added to refs/heads/main by this push:
new 9f241d42332 Add toolset overview and MCP tool filtering example to
`common.ai` docs (#74378)
9f241d42332 is described below
commit 9f241d4233248ee85a0f5588965114fac7539a1b
Author: Kaxil Naik <[email protected]>
AuthorDate: Wed Oct 7 14:15:50 2026 +0100
Add toolset overview and MCP tool filtering example to `common.ai` docs
(#74378)
The toolsets index now opens with a table of every toolset, what the agent
gets from it and the parameters that limit it, plus the two controls that
work
on any toolset (`.approval_required()` and `.filtered()`). The MCP guide
gains
a runnable `.filtered()` example with the model-facing refusal captured
from a
real run, and the SQL, hook and MCP guides link to per-tool approval with
its
once-per-Dag-run limit.
Smaller fixes: ObjectStorageToolset rows in the credentials table and the
barriers section, duplicate snippets removed from the DataFusion and Agent
Skills guides, a stale pinned-argument sentence, an undefined logger in the
LoggingToolset snippet, and an internal connection name in example_mcp.py.
---
providers/common/ai/docs/toolsets/datafusion.rst | 26 +--
providers/common/ai/docs/toolsets/hook.rst | 3 +
providers/common/ai/docs/toolsets/index.rst | 207 ++++++++++++---------
providers/common/ai/docs/toolsets/logging.rst | 9 +-
providers/common/ai/docs/toolsets/mcp.rst | 45 ++++-
.../common/ai/docs/toolsets/object_storage.rst | 4 +-
providers/common/ai/docs/toolsets/skills.rst | 15 +-
providers/common/ai/docs/toolsets/sql.rst | 7 +-
.../common/ai/example_dags/example_mcp.py | 35 +++-
9 files changed, 211 insertions(+), 140 deletions(-)
diff --git a/providers/common/ai/docs/toolsets/datafusion.rst
b/providers/common/ai/docs/toolsets/datafusion.rst
index a1aa337b38c..31acfae273d 100644
--- a/providers/common/ai/docs/toolsets/datafusion.rst
+++ b/providers/common/ai/docs/toolsets/datafusion.rst
@@ -175,27 +175,11 @@ Iceberg is looked up by ``db_name`` instead, and
``DataSourceConfig`` raises
error message against regular expressions, which a wording change upstream
can
quietly defeat.
-**A real example.** The same bucket as the ``HookToolset`` example on
:doc:`hook`, reached
-the other way. Rather than exposing ``list_keys`` and ``read_key`` and leaving
-the agent to reassemble files, this registers the prefix as a table and lets it
-write SQL:
-
-.. code-block:: python
-
- from airflow.providers.common.ai.toolsets.datafusion import
DataFusionToolset
- from airflow.providers.common.sql.config import DataSourceConfig
-
- toolset = DataFusionToolset(
- datasource_configs=[
- DataSourceConfig(
- conn_id="aws_default",
- table_name="sales",
- uri="s3://my-bucket/data/sales/",
- format="parquet",
- ),
- ],
- max_rows=100,
- )
+**Compared with the hook route.** The ``sales`` table at the top of this page
+reaches an S3 prefix the other way from the ``HookToolset`` example on
+:doc:`hook`. Rather than exposing
+``list_keys`` and ``read_key`` and leaving the agent to reassemble files, it
+registers the prefix as a table and lets the agent write SQL.
Which of the two fits depends on the question. "Read me this object" is a hook
method. "What were last quarter's returns by region" is a query, and expressing
diff --git a/providers/common/ai/docs/toolsets/hook.rst
b/providers/common/ai/docs/toolsets/hook.rst
index 84d93871000..89df4662326 100644
--- a/providers/common/ai/docs/toolsets/hook.rst
+++ b/providers/common/ai/docs/toolsets/hook.rst
@@ -206,6 +206,9 @@ reflection-based adapter, so the work is choosing the
method list.
``read_key`` is exposed, the agent picks the key within the pinned bucket;
the
:ref:`defense-layer table <toolset-defense-layers>` states this outright.
Choose
methods whose worst case you accept, not methods you intend to constrain
later.
+ To expose a method that changes something and have a person approve the call
+ first, wrap the toolset with ``.approval_required()``. A task instance can
pause
+ for approval once per Dag run; see :doc:`../tool_approval`.
- Its calls act as barriers. The tools are registered with ``sequential=True``
and each hook method runs in a worker thread, one blocking hook call at a
time
in the task process, so a slow call holds up every other tool the model
emitted
diff --git a/providers/common/ai/docs/toolsets/index.rst
b/providers/common/ai/docs/toolsets/index.rst
index eb9c0525448..d4a87b5ce85 100644
--- a/providers/common/ai/docs/toolsets/index.rst
+++ b/providers/common/ai/docs/toolsets/index.rst
@@ -20,27 +20,110 @@
Toolsets
========
+A toolset is a group of tools an agent may call. Pass toolsets to
+:class:`~airflow.providers.common.ai.operators.agent.AgentOperator` or
+``@task.agent`` in ``toolsets=[...]``. The model sees each tool's name,
description
+and arguments, and what each call returns. The connection a toolset
authenticates
+with is resolved in the worker and is not part of any of that.
+
+.. list-table::
+ :widths: 22 40 38
+ :header-rows: 1
+
+ * - Toolset
+ - What the agent gets
+ - What you limit it with
+ * - :doc:`HookToolset <hook>`
+ - The methods you list from any Airflow hook, one tool per method
+ - ``allowed_methods`` (required), ``pinned_arguments``
+ * - :doc:`SQLToolset <sql>`
+ - ``list_tables``, ``get_schema``, ``query`` and ``check_query`` against
a DBAPI
+ database
+ - ``allowed_tables``, ``allow_writes`` (off by default), ``max_rows``,
+ ``max_result_bytes``
+ * - :doc:`ObjectStorageToolset <object_storage>`
+ - ``list_files``, ``get_file_info`` and ``read_file`` under one
object-storage path.
+ It cannot write.
+ - ``path``, the root every requested path is checked against;
+ ``max_read_bytes``, ``max_output_bytes``
+ * - :doc:`DataFusionToolset <datafusion>`
+ - SQL over Parquet, CSV and Avro files and Iceberg tables, run in the
worker
+ - The tables you register in ``datasource_configs``, ``allow_writes``
(off by
+ default), ``max_rows``
+ * - :doc:`MCPToolset <mcp>`
+ - Every tool the MCP server behind ``mcp_conn_id`` exposes
+ - ``.filtered()`` to offer only some of them
+ * - :doc:`AgentSkillsToolset <skills>`
+ - ``SKILL.md`` instruction bundles that the model loads when it needs one
+ - ``exclude_tools={"run_skill_script"}`` to stop skill scripts running on
the
+ worker; ``exclude_resources``
+ * - :doc:`SandboxToolset <../sandbox/index>`
+ - A shell and a filesystem in a sandbox isolated from the worker process
+ - ``SandboxSpec`` (``block_network`` is on by default,
``allow_egress_to``),
+ command timeouts
+ * - :doc:`ManagedAgentToolset <managed_agent>`
+ - One tool that sends a prompt to an agent running on a cloud vendor's
+ infrastructure
+ - ``timeout`` for each call; what the remote agent may touch is set at the
+ vendor
+
+Two controls work on every toolset. ``.approval_required()`` pauses the task
before a
+matching call runs, until a person approves or rejects it on the **Required
Actions**
+page. It needs Airflow 3.3 or later, pauses a task instance at most once per
Dag run,
+and does not combine with ``durable=True`` or a ``SandboxToolset`` that
provisions its
+own sandbox (:doc:`../tool_approval` lists every limit). ``.filtered()`` drops
tools
+from what the model is offered, as the :ref:`MCP guide
<howto/toolset:mcp-filtered>`
+shows.
+
+:doc:`logging` and :doc:`langchain` do not reach a system of their own: one
wraps a
+toolset to log its calls, the other converts a toolset for a LangChain agent.
+
+.. toctree::
+ :titlesonly:
+ :hidden:
+
+ Airflow hooks as tools <hook>
+ SQL databases <sql>
+ Files on object storage <object_storage>
+ Files with DataFusion <datafusion>
+ MCP servers <mcp>
+ Agent Skills <skills>
+ Sandboxed execution <../sandbox/index>
+ Vendor-managed agents <managed_agent>
+ LangChain tools <langchain>
+ Tool call logging <logging>
+
+Importing toolsets
+------------------
+
+Six toolsets import from the ``airflow.providers.common.ai.toolsets`` package
root:
+``HookToolset``, ``SQLToolset``, ``ObjectStorageToolset``, ``MCPToolset``,
+``SandboxToolset`` and ``ManagedAgentToolset``. Import the other three from
their own
+submodules::
+
+ from airflow.providers.common.ai.toolsets.datafusion import
DataFusionToolset
+ from airflow.providers.common.ai.toolsets.logging import LoggingToolset
+ from airflow.providers.common.ai.toolsets.skills import AgentSkillsToolset
+
+Every toolset here implements pydantic-ai's
+`AbstractToolset <https://ai.pydantic.dev/toolsets/>`__ interface, so it works
in any
+pydantic-ai ``Agent``. ``AgentOperator`` accepts any ``AbstractToolset`` too,
such as
+pydantic-ai's own ``MCPToolset`` or a third-party toolset; those miss the
connection
+handling the toolsets on this page add.
+
Choosing a toolset
------------------
-Each toolset's guide documents how to configure it. This section answers the
-question that comes before that one: you have a system you want an agent to
-reach, so which route do you take, and what does each route give up?
-
-Read the table below by what you already have, not by what a toolset is called.
-When two routes both work, the deciding factor is rarely what each one can do.
-It is what each one cannot do, and every route has a short list.
-
-More than one row can be true at once, and the rows are not exclusive: one
agent
-can carry several toolsets. Two questions break the ties. *Whose credential is
-it?* Prefer the route whose credential is an Airflow connection somebody on
-your side already reviewed. *Whose tool list is it?* Prefer the route whose
-exposed surface you chose rather than inherited. The pair that most often
-overlaps is an Airflow hook and a vendor MCP server reaching the same target;
-both questions point at the hook, because its credential is the connection and
-``allowed_methods`` is a list you write. Reach for the server when its tools
-cover work the hook does not expose, or when the alternative is re-wrapping
that
-API by hand.
+Pick a route by what you already have. More than one row of the table below
can be
+true at once, and the rows are not exclusive: one agent can carry several
toolsets.
+Two questions break the ties. *Whose credential is it?* Prefer the route whose
+credential is an Airflow connection somebody on your side already reviewed.
*Whose
+tool list is it?* Prefer the route whose exposed surface you chose rather than
+inherited. The pair that most often overlaps is an Airflow hook and a vendor
MCP
+server reaching the same target; both questions point at the hook, because its
+credential is the connection and ``allowed_methods`` is a list you write.
Reach for
+the server when its tools cover work the hook does not expose, or when the
+alternative is re-wrapping that API by hand.
Those two questions do not separate ``HookToolset`` from ``SQLToolset`` when
the
target is a DBAPI database, because both answer them the same way. A third one
@@ -86,74 +169,6 @@ The hook, SQL, object storage, DataFusion, MCP, Agent
Skills and managed-agent g
example that exists in this repository, and where its credentials and its work
come
from. :doc:`../sandbox/index` carries the same section for ``SandboxToolset``.
-Toolset guides
---------------
-
-.. toctree::
- :titlesonly:
-
- Airflow hooks as tools <hook>
- SQL databases <sql>
- Files on object storage <object_storage>
- Files with DataFusion <datafusion>
- MCP servers <mcp>
- Agent Skills <skills>
- Sandboxed execution <../sandbox/index>
- Vendor-managed agents <managed_agent>
- LangChain tools <langchain>
- Tool call logging <logging>
-
-The toolsets
-------------
-
-Airflow's 350+ provider hooks already have typed methods, rich docstrings,
-and managed credentials. Toolsets expose them as pydantic-ai tools so that
-LLM agents can call them during multi-turn reasoning.
-
-Six toolsets are exported directly from the
``airflow.providers.common.ai.toolsets``
-package root:
-
-- :class:`~airflow.providers.common.ai.toolsets.hook.HookToolset`: generic
- adapter for any Airflow Hook. Guide: :doc:`hook`.
-- :class:`~airflow.providers.common.ai.toolsets.sql.SQLToolset`: curated
- 4-tool database toolset. Guide: :doc:`sql`.
--
:class:`~airflow.providers.common.ai.toolsets.object_storage.ObjectStorageToolset`:
- read-only access to the files under one object-storage path. Guide:
- :doc:`object_storage`.
-- :class:`~airflow.providers.common.ai.toolsets.mcp.MCPToolset`: connect to
- `MCP servers <https://modelcontextprotocol.io/>`__ configured via Airflow
- connections. Guide: :doc:`mcp`.
-- :class:`~airflow.providers.common.ai.toolsets.sandbox.SandboxToolset`: give
- the agent a shell and a filesystem inside an isolated sandbox, off the
- Airflow worker. Guide: :doc:`../sandbox/index`.
--
:class:`~airflow.providers.common.ai.toolsets.managed_agent.ManagedAgentToolset`:
- expose a **vendor-managed agent**, one whose reasoning loop runs on a cloud
- provider's infrastructure, as a single tool. Takes any
- :class:`~airflow.providers.common.ai.managed_agents.base.ManagedAgentClient`,
- which the Amazon and Google providers produce from their hooks and which
-
:class:`~airflow.providers.common.ai.managed_agents.failover.FailoverManagedAgentClient`
- composes across clouds. Guide: :doc:`managed_agent`.
-
-Three more toolsets (:doc:`datafusion`, :doc:`logging`, :doc:`skills`) are not
-re-exported from the package root, so import each of them from its own
submodule::
-
- from airflow.providers.common.ai.toolsets.datafusion import
DataFusionToolset
- from airflow.providers.common.ai.toolsets.logging import LoggingToolset
- from airflow.providers.common.ai.toolsets.skills import AgentSkillsToolset
-
-All of these toolsets implement pydantic-ai's
-`AbstractToolset <https://ai.pydantic.dev/toolsets/>`__ interface and can be
-passed to any pydantic-ai ``Agent``, including via
-:class:`~airflow.providers.common.ai.operators.agent.AgentOperator`.
-
-.. note::
-
- ``AgentOperator`` accepts **any** ``AbstractToolset`` implementation, not
- just the Airflow-native toolsets above. pydantic-ai's own ``MCPToolset``
- (built over a FastMCP transport) and third-party toolsets work too. The
- Airflow-native toolsets add connection management, secret backend
- integration, and the connection UI, but you are not locked in.
-
Where the credentials come from
-------------------------------
@@ -176,6 +191,10 @@ making on purpose rather than inheriting.
* - ``SQLToolset``
- ``db_conn_id``, via ``BaseHook.get_connection``
- Worker process, against the database
+ * - ``ObjectStorageToolset``
+ - ``conn_id``; with ``conn_id=None``, the store's default credentials,
such as
+ the worker's AWS role
+ - Worker process, against the store
* - ``DataFusionToolset``
- ``conn_id`` on each ``DataSourceConfig``
- Worker process (embedded engine)
@@ -210,18 +229,20 @@ right companion when you do.
Tool calls as barriers
----------------------
-One behaviour cuts across these routes rather than telling them apart.
``HookToolset``, ``SQLToolset``, ``DataFusionToolset`` and ``SandboxToolset``
each build their own tool definitions and set ``sequential=True`` on them,
which
pydantic-ai treats as a barrier: the tool runs alone, tools the model emitted
before it finish first, and tools emitted after it start only once it returns.
A slow call on any of those four therefore holds up the rest of that step, not
-just its own toolset. ``ManagedAgentToolset`` sets ``sequential=False``
-deliberately, because the wait it introduces is remote. ``MCPToolset`` and
-``AgentSkillsToolset`` define no tools of their own: they pass through whatever
-the upstream toolset declares, so the setting is not theirs to make. Do not
read
-this as a reason to choose one route over another; read it as something to
expect
-from all four.
+just its own toolset. Expect this from all four; it is not a reason to prefer
+one of them. ``ManagedAgentToolset`` sets ``sequential=False``
+deliberately, because the wait it introduces is remote.
``ObjectStorageToolset``
+leaves it unset, so its calls are not barriers, but each of them still waits
for a
+hook, SQL or DataFusion call in progress: the hook, SQL, DataFusion and object
storage
+toolsets run their blocking work in worker threads under one lock per task
process.
+``MCPToolset`` and ``AgentSkillsToolset`` define no tools of their own: they
pass
+through whatever the upstream toolset declares, so the setting is not theirs to
+make.
.. _toolset-retry-budget:
diff --git a/providers/common/ai/docs/toolsets/logging.rst
b/providers/common/ai/docs/toolsets/logging.rst
index 816df3a217a..76e300bdbe4 100644
--- a/providers/common/ai/docs/toolsets/logging.rst
+++ b/providers/common/ai/docs/toolsets/logging.rst
@@ -34,8 +34,9 @@ in real time. ``AgentOperator`` applies it automatically (see
from airflow.providers.common.ai.toolsets.logging import LoggingToolset
from airflow.providers.common.ai.toolsets.sql import SQLToolset
- sql_toolset = SQLToolset(db_conn_id="my_db")
- logged_toolset = LoggingToolset(wrapped=sql_toolset, logger=my_logger)
+ logged_toolset = LoggingToolset(wrapped=SQLToolset(db_conn_id="my_db"))
-Each tool call produces two INFO log lines (name + timing) and optional
-DEBUG-level argument logging. Exceptions are logged and re-raised.
+Each call logs the tool's name and how long it took at INFO, inside a
collapsible
+``::group::`` block in the task log, and its arguments at DEBUG. A call that
raises
+is logged with its traceback and the exception is re-raised. Pass ``logger`` to
+send the lines to a logger other than the toolset module's own.
diff --git a/providers/common/ai/docs/toolsets/mcp.rst
b/providers/common/ai/docs/toolsets/mcp.rst
index 639b5a665e1..64b69d3723f 100644
--- a/providers/common/ai/docs/toolsets/mcp.rst
+++ b/providers/common/ai/docs/toolsets/mcp.rst
@@ -166,6 +166,39 @@ Using multiple MCP servers
],
)
+.. _howto/toolset:mcp-filtered:
+
+Offer only some of a server's tools
+-----------------------------------
+
+``MCPToolset`` offers the model every tool the server lists. To offer fewer,
wrap it
+with ``.filtered()``. The function you pass receives the run context and each
tool's
+definition, and the model is offered only the tools it returns ``True`` for:
+
+.. exampleinclude::
/../../ai/src/airflow/providers/common/ai/example_dags/example_mcp.py
+ :language: python
+ :start-after: [START howto_toolset_mcp_filtered]
+ :end-before: [END howto_toolset_mcp_filtered]
+
+Against a server that lists ``get_forecast``, ``get_alerts`` and
``delete_alert``, the
+model is offered the first two. If it calls ``delete_alert`` anyway, the call
never
+reaches the server, and the model gets this back instead of a result:
+
+.. code-block:: text
+
+ Unknown tool name: 'delete_alert'. Available tools: 'get_alerts',
'get_forecast'
+
+ Fix the errors and try again.
+
+With ``tool_prefix`` set, the filter sees the prefixed names, such as
+``weather_get_forecast``. The filter runs in the worker and changes only what
this
+agent is offered. It does not change what the server lets the connection's
credential
+do, so give the connection the narrowest credential the server accepts.
+
+To keep a tool available but have a person approve a call to it before it
runs, wrap
+the toolset with ``.approval_required()`` instead. A task instance can pause
for
+approval once per Dag run; see :doc:`../tool_approval`.
+
Direct pydantic-ai MCP toolsets
-------------------------------
@@ -209,12 +242,12 @@ yourself.
``MCPToolset`` is itself built on the same ``AbstractToolset`` base every
toolset in this provider extends, a Dag author can call ``.filtered()`` to
subset the advertised tool list using a filter function that inspects each
- tool's definition. That filtering happens on the client: it narrows what
- the agent is offered, it does not revoke or authorize anything on the
- server, and unlike ``allowed_methods`` on ``HookToolset``, which is required
- and rejects an empty list, nothing here requires you to set a filter. The
- defense-layer table is explicit that a server can expose shell, filesystem
- or network access.
+ tool's definition (see :ref:`howto/toolset:mcp-filtered`). That filtering
+ happens on the client: it narrows what the agent is offered, it does not
+ revoke or authorize anything on the server, and unlike ``allowed_methods`` on
+ ``HookToolset``, which is required and rejects an empty list, nothing here
+ requires you to set a filter. The defense-layer table is explicit that a
+ server can expose shell, filesystem or network access.
- It cannot guarantee the credential came from a connection. ``mcp_conn_id`` is
the default path, but ``token_provider`` and ``env_provider`` are your own
callables and are free to read an environment variable, a file, or an
entirely
diff --git a/providers/common/ai/docs/toolsets/object_storage.rst
b/providers/common/ai/docs/toolsets/object_storage.rst
index ed3d6577724..6a7a0b8c5f6 100644
--- a/providers/common/ai/docs/toolsets/object_storage.rst
+++ b/providers/common/ai/docs/toolsets/object_storage.rst
@@ -142,7 +142,9 @@ storage already uses; Parquet and Avro need this provider's
``parquet`` or ``avr
extra. To hand the model a known file rather than let it find one,
``@task.llm_file_analysis``
reads the file for it. For read-only access through a hook's own methods,
``HookToolset(S3Hook(), allowed_methods=["list_keys", "read_key"])`` works
too, but the
-model then chooses the bucket and key, where this toolset holds it under one
root.
+model then chooses the bucket and key. Pinning ``bucket_name`` with
``pinned_arguments``
+fixes the bucket and still leaves the model any key in it, where this toolset
keeps the
+model under one root.
**What it cannot do**
diff --git a/providers/common/ai/docs/toolsets/skills.rst
b/providers/common/ai/docs/toolsets/skills.rst
index 1e0b1ed6dd7..0b0bc9d42f8 100644
--- a/providers/common/ai/docs/toolsets/skills.rst
+++ b/providers/common/ai/docs/toolsets/skills.rst
@@ -199,17 +199,10 @@ discoverable. See :ref:`agent-skills` for the layout.
that clone once per run. A local directory is read in place and
costs nothing.
-**A real example.** ``example_agent_skills.py`` loads skills from a local
-directory:
-
-.. exampleinclude::
/../../ai/src/airflow/providers/common/ai/example_dags/example_agent_skills.py
- :language: python
- :start-after: [START howto_operator_agent_skills_local]
- :end-before: [END howto_operator_agent_skills_local]
-
-The two skills it ships, ``aip-tracker`` and ``sql-reporting``, are procedural
by
-nature: neither adds an endpoint the agent could not already reach. That is the
-signal you are on the right route.
+**A real example.** The local-directory example at the top of this page comes
+from ``example_agent_skills.py``. The two skills it ships, ``aip-tracker`` and
+``sql-reporting``, are procedural by nature: neither adds an endpoint the agent
+could not already reach. That is the signal you are on the right route.
**Credentials and where it runs.** A local directory needs no credential. A
private repository goes through
diff --git a/providers/common/ai/docs/toolsets/sql.rst
b/providers/common/ai/docs/toolsets/sql.rst
index 6d0a85b6d71..51a8d602071 100644
--- a/providers/common/ai/docs/toolsets/sql.rst
+++ b/providers/common/ai/docs/toolsets/sql.rst
@@ -297,7 +297,12 @@ Parameters
introspection. Schema-qualified ``allowed_tables`` entries override it per
table.
- ``allow_writes``: Allow data-modifying SQL (INSERT, UPDATE, DELETE, etc.).
Default ``False`` -- only SELECT-family and read-only metadata
- (``DESCRIBE``/``SHOW``) statements are permitted.
+ (``DESCRIBE``/``SHOW``) statements are permitted. To have a person approve a
+ ``query`` call before it runs, wrap the toolset with
``.approval_required()`` and
+ return ``tool_def.name == "query"`` from its function. Reads and writes both
go
+ through ``query``, and a task instance can pause for approval once per Dag
run, so
+ this fits an agent that runs one query, such as a single write; see
+ :doc:`../tool_approval`.
- ``max_rows``: Maximum rows returned from the ``query`` tool. Default ``50``.
Rows beyond it are not read out of a DBAPI cursor; what the driver has
already
transferred is its own call. See :ref:`bounded-query-results`.
diff --git
a/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_mcp.py
b/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_mcp.py
index 023360b3516..e2090b752dc 100644
---
a/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_mcp.py
+++
b/providers/common/ai/src/airflow/providers/common/ai/example_dags/example_mcp.py
@@ -74,7 +74,36 @@ example_mcp_multiple_servers()
# ---------------------------------------------------------------------------
-# 3. Direct PydanticAI MCP toolsets (no Airflow connection needed)
+# 3. Offer the agent only some of a server's tools
+# ---------------------------------------------------------------------------
+
+
+# [START howto_toolset_mcp_filtered]
+READ_ONLY_WEATHER_TOOLS = {"get_forecast", "get_alerts"}
+
+
+@dag(tags=["example"])
+def example_mcp_filtered_tools():
+ """Offer the agent only the MCP tools the Dag author listed."""
+ AgentOperator(
+ task_id="forecast_agent",
+ prompt="Will it rain in London tomorrow?",
+ llm_conn_id="pydanticai_default",
+ toolsets=[
+ MCPToolset(mcp_conn_id="weather_mcp").filtered(
+ lambda ctx, tool_def: tool_def.name in READ_ONLY_WEATHER_TOOLS
+ ),
+ ],
+ )
+
+
+# [END howto_toolset_mcp_filtered]
+
+example_mcp_filtered_tools()
+
+
+# ---------------------------------------------------------------------------
+# 4. Direct PydanticAI MCP toolsets (no Airflow connection needed)
# ---------------------------------------------------------------------------
# AgentOperator accepts any PydanticAI AbstractToolset, including MCPToolset
# directly. Use this for prototyping or when you want full PydanticAI control.
@@ -94,7 +123,7 @@ example_mcp_multiple_servers()
# ---------------------------------------------------------------------------
-# 4. Stdio server with a minted secret in the subprocess environment
+# 5. Stdio server with a minted secret in the subprocess environment
# ---------------------------------------------------------------------------
# For local stdio MCP servers that read credentials from their own environment
# (e.g. a server that needs a Splunk API key), pass env_provider instead of
@@ -124,7 +153,7 @@ def example_mcp_stdio_env_provider():
llm_conn_id="pydanticai_default",
system_prompt="You are a support triage agent with access to MCP
tools.",
toolsets=[
- MCPToolset(mcp_conn_id="spacefarer_mcp",
env_provider=_mint_splunk_env),
+ MCPToolset(mcp_conn_id="splunk_tools_mcp",
env_provider=_mint_splunk_env),
],
)