kaxil opened a new pull request, #73994: URL: https://github.com/apache/airflow/pull/73994
An agent re-sends its tool definitions, system prompt and conversation on every request. A tool-calling agent makes several requests per task, and a mapped `@task.agent` repeats all of them per map index with the same system prompt. OpenAI and Gemini cache that repeated prefix on their own; Anthropic, Bedrock and OpenRouter cache only when each request is marked, and pydantic-ai spells the marks differently per provider (`anthropic_cache_*`, `bedrock_cache_*`, `openrouter_cache_*` in `model_settings`). In practice nobody sets them, so these agents pay full input price for the same prefix over and over. This adds `cache_prompt: bool = True` to `AgentOperator` (and so `@task.agent`). It turns on each provider's own prompt caching for the tool definitions, the system prompt and the latest message, and is a no-op for providers that cache automatically. Setting it `False` restores the old requests exactly. ## Evidence The same Dag against `claude-opus-4-8` (and one task on `gpt-5.4-mini`), 60 lines of triage policy as the system prompt (about 6,000 tokens) and one tool call per task, so two requests each: | Task | Input tokens | Cache read | Cache write | |---|---:|---:|---:| | `cache_prompt=False` | 12,160 | 0 | 0 | | First task, cold cache | 12,167 | 6,027 | 6,136 | | Mapped task, each of 3 map indexes | ~12,150 | ~11,940 | ~200 | | Caller sets `anthropic_cache=True` itself | 12,143 | 12,139 | 0 | | OpenAI model, same flag | 7,890 | 3,584 (OpenAI's own) | 0 | The first request writes the prefix and the second reads it back. Every map index then reads the tools and system prompt the first task wrote. On a second run of the same Dag, pydantic-ai priced the task with caching off at **$0.065** and the same task with caching on and a warm cache at **$0.011**. The task log shows the split under the run summary:    ## Design rationale **A pydantic-ai capability, not a settings merge in the operator.** `PromptCaching.get_model_settings()` returns a callable, which pydantic-ai calls with the settings merged so far: the model's own, the agent's `model_settings` (static or callable), and a spec file's. It fills in only the providers none of those configure. Merging a dict into `agent_params["model_settings"]` could not see a spec file or a callable, and would have had to guess who wins. **Anything the caller set for a provider takes that provider over entirely.** Setting any `anthropic_cache*` key means nothing is added for Anthropic, rather than filling in the keys that are missing. pydantic-ai refuses `anthropic_cache` together with `anthropic_cache_messages`, so filling keys one by one would turn a valid caller setting into an error. A `CachePoint` in the prompt or the message history turns all of it off, because Anthropic requires a longer-lived cache entry to come before a shorter one, and our 5-minute marks on the tools and system prompt would land ahead of a caller's 1-hour mark. **`anthropic_cache_messages` rather than the top-level `anthropic_cache`.** Both move the cache mark to the latest message on each request. pydantic-ai documents the per-block form as the one that works with Anthropic-compatible gateways that lack the top-level parameter, and on Bedrock and Vertex `anthropic_cache` falls back to it anyway. A caller who prefers `anthropic_cache` sets it and gets it, as the table shows. **No provider detection.** Each pydantic-ai model reads only its own provider's settings and ignores the rest, and Bedrock and OpenRouter add the marks only for models whose profile supports caching. One set of settings therefore covers a `FallbackModel` chain that spans providers, with nothing resolved in the hook. **Why on by default.** The people who save the most (long system prompts, tool loops, fan-out over rows) are the ones least likely to find three provider-specific flags. A prompt shorter than Anthropic's minimum length for caching (512 to 4,096 tokens depending on the model) is not cached and costs nothing extra. The docs name the cases where caching costs more than it saves: a single long request never repeated within five minutes, a final tool result much larger than the prompt, and map indexes that all start at the same moment. Cache settings are excluded from the durable-execution request fingerprint, so switching `cache_prompt` between attempts does not re-run steps an earlier attempt completed. ## Gotchas - `input_tokens` in the log and the `usage` XCom includes the cached tokens; the new `LLM prompt cache:` line is where the split shows. The XCom does not carry the split. - Only `AgentOperator` and `@task.agent` get the flag. `LLMOperator` and the other `LLM*` operators build agents the same way and could take it in a follow-up. --- * Read the **[Pull Request Guidelines](https://github.com/apache/airflow/blob/main/contributing-docs/05_pull_requests.rst#pull-request-guidelines)** for more information. Note: commit author/co-author name and email in commits become permanently public when merged. * For fundamental code changes, an Airflow Improvement Proposal ([AIP](https://cwiki.apache.org/confluence/display/AIRFLOW/Airflow+Improvement+Proposals)) is needed. * When adding dependency, check compliance with the [ASF 3rd Party License Policy](https://www.apache.org/legal/resolved.html#category-x). * For significant user-facing changes create newsfragment: `{pr_number}.significant.rst`, in [airflow-core/newsfragments](https://github.com/apache/airflow/tree/main/airflow-core/newsfragments). You can add this file in a follow-up commit after the PR is created so you know the PR number. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
