kaxil opened a new pull request, #73957:
URL: https://github.com/apache/airflow/pull/73957

   `SQLToolset`, `HookToolset`, `DataFusionToolset` and `ObjectStorageToolset` 
pinned `max_retries=1` on every tool, which overrides the agent's `retries`. An 
agent given `retries=3` still ended its run on the second failed query in a 
row, so a model that needed two tries to get a column name right could not 
recover. The only workaround was to subclass the toolset and rewrite 
`max_retries` in `get_tools`.
   
   Each toolset now takes `max_retries`, defaulting to `None`, which uses the 
run's own tool retry budget (`ctx.max_retries`, the agent's `retries`). That is 
the order pydantic-ai's own `FunctionToolset` and `MCPServer` use, so these 
toolsets now behave like pydantic-ai's. The fallback lives once on 
`AirflowToolset`.
   
   **Agents that already set `retries` now apply it to these toolsets.** They 
used to allow exactly one correction whatever `retries` said. `retries=0` now 
fails the run on the first bad query, and a large integer `retries` meant for 
output validation also lets a failing database be queried that many times. A 
dict such as `retries={"output": 3}` raises only the output budget, and 
`max_retries=1` on a toolset keeps the old behaviour. The new docs section 
spells this out.
   
   **What the budget counts differs by toolset**, and the docs now say so. 
`SQLToolset` and `DataFusionToolset` turn every query error into `ModelRetry`. 
`HookToolset` counts invalid arguments and attempts to change a pinned 
argument; an exception from the hook fails the run straight away. 
`ObjectStorageToolset` counts invalid arguments only, since a failed read comes 
back as `ToolFailed`, which pydantic-ai does not count.
   
   **Outside a pydantic-ai agent** (the LangChain, Strands and Google ADK 
bridges) there is no agent budget, so the bridge's placeholder `RunContext` now 
carries pydantic-ai's default of one correction. Each bridged call also gets 
its own copy with the tool's retry count, as pydantic-ai's `ToolManager` does, 
so `ctx.last_attempt` means the same through a bridge as inside a run. One side 
effect: a bare `FunctionToolset` or `MCPToolset` bridged without its own 
`max_retries` now gets one correction instead of none, matching a default 
pydantic-ai agent.
   
   `SandboxToolset` still pins `max_retries=2`. Moving it to the agent's budget 
would lower its default to one, so it is left for a follow-up.
   
   Run end to end on a local Airflow 3.4.0 with a real `SQLToolset` against 
SQLite. The model is a scripted pydantic-ai `FunctionModel` that misspells a 
column twice before reading the error, since a real model will not make the 
same mistake on demand. With `agent_params={"retries": 3}`, the first two 
queries fail with `no such column: totl`, the third uses `total`, and the task 
succeeds:
   
   ![Agent retries 3: two failed queries, then the corrected query 
succeeds](./afl267-agent_retries_3.png)
   
   The same Dag with the default budget still fails on the second error, as 
before:
   
   ![Default budget: the second failed query ends the 
run](./afl267-default_retry_budget.png)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to