kaxil commented on code in PR #71575:
URL: https://github.com/apache/airflow/pull/71575#discussion_r4127938893


##########
providers/common/ai/src/airflow/providers/common/ai/durable/fingerprint.py:
##########
@@ -98,19 +123,107 @@ def _strip_volatile(messages_dump: list[dict[str, Any]]) 
-> list[dict[str, Any]]
     return stripped
 
 
+def _is_dataclass_instance(value: Any) -> bool:
+    return dataclasses.is_dataclass(value) and not isinstance(value, type)
+
+
+def _refuse_iterators(value: Any) -> None:
+    """
+    Raise ``TypeError`` if ``value`` holds an iterator anywhere inside it.
+
+    Rendering an iterator consumes it, and pydantic validates an 
``Iterable[T]``
+    tool parameter lazily into a ``ValidatorIterator``. Tool arguments are
+    fingerprinted before the tool runs, so rendering them would drain the 
tool's
+    own input: the tool would see an empty sequence, and that wrong result 
would be
+    cached under the fingerprint of the full one. Refusing degrades the step 
to the
+    not-cached path instead.
+    """
+    if value is None or isinstance(value, (str, bytes, bytearray, bool, int, 
float)):
+        return
+    if isinstance(value, Iterator):
+        raise TypeError(f"cannot fingerprint {type(value).__name__} without 
consuming it")
+    if isinstance(value, Mapping):
+        children: Iterable[Any] = value.values()
+    elif isinstance(value, (list, tuple, set, frozenset)):
+        children = value
+    elif isinstance(value, BaseModel):
+        children = [getattr(value, name, None) for name in 
type(value).model_fields]

Review Comment:
   This walks `model_fields` only, so a model's extras never reach the guard, 
but `to_jsonable_python` does render them. A tool argument typed as a model 
with `extra="allow"` and `__pydantic_extra__: dict[str, Iterable[int]]` 
validates each extra into a `ValidatorIterator`, even from JSON args, and the 
render drains it before the tool runs. I ran it through a real `Agent` with 
`FunctionModel` and `FunctionToolset`, calling a tool that sums the extra: the 
fingerprint came back non-`None` and the tool returned 0 for `[1, 2, 3]`. That 
is the drain-then-cache-the-wrong-result path the `Iterable[T]` refusal exists 
for. Adding `(value.model_extra or {}).values()` to the children here closes 
it: the same argument then fingerprints as `None` and the extra still sums to 6.



##########
providers/common/ai/src/airflow/providers/common/ai/durable/fingerprint.py:
##########
@@ -98,19 +123,107 @@ def _strip_volatile(messages_dump: list[dict[str, Any]]) 
-> list[dict[str, Any]]
     return stripped
 
 
+def _is_dataclass_instance(value: Any) -> bool:
+    return dataclasses.is_dataclass(value) and not isinstance(value, type)
+
+
+def _refuse_iterators(value: Any) -> None:
+    """
+    Raise ``TypeError`` if ``value`` holds an iterator anywhere inside it.
+
+    Rendering an iterator consumes it, and pydantic validates an 
``Iterable[T]``
+    tool parameter lazily into a ``ValidatorIterator``. Tool arguments are
+    fingerprinted before the tool runs, so rendering them would drain the 
tool's
+    own input: the tool would see an empty sequence, and that wrong result 
would be
+    cached under the fingerprint of the full one. Refusing degrades the step 
to the
+    not-cached path instead.
+    """
+    if value is None or isinstance(value, (str, bytes, bytearray, bool, int, 
float)):
+        return
+    if isinstance(value, Iterator):
+        raise TypeError(f"cannot fingerprint {type(value).__name__} without 
consuming it")
+    if isinstance(value, Mapping):
+        children: Iterable[Any] = value.values()
+    elif isinstance(value, (list, tuple, set, frozenset)):
+        children = value
+    elif isinstance(value, BaseModel):
+        children = [getattr(value, name, None) for name in 
type(value).model_fields]
+    elif _is_dataclass_instance(value):
+        children = [getattr(value, field.name, None) for field in 
dataclasses.fields(value)]
+    else:
+        return
+    for child in children:
+        _refuse_iterators(child)
+
+
+def _order_sets(value: Any, rendered: Any) -> Any:
+    """
+    Return ``rendered`` with every list that pydantic rendered from a set 
sorted.
+
+    ``rendered`` is pydantic's JSON rendering of ``value`` and is otherwise 
kept as
+    is, so pydantic stays the only renderer: bytes, dates, dict keys, ``NaN`` 
and
+    custom serializers come out exactly as a json-mode dump renders them. The 
one
+    thing pydantic cannot do stably is order a set. It lists the members in
+    iteration order, which for strings follows the interpreter's hash seed, and
+    every task attempt is a fresh process, so a ``set[str]`` would hash 
differently
+    on each attempt and never replay. ``value`` is walked alongside 
``rendered``
+    only to find those lists; a branch whose rendering does not line up with 
the
+    object, such as one with a custom serializer, is left exactly as rendered.
+    """
+    if isinstance(value, (set, frozenset)):
+        if not isinstance(rendered, list) or len(rendered) != len(value):
+            return rendered
+        members = [_order_sets(member, item) for member, item in zip(value, 
rendered)]
+        return sorted(members, key=lambda member: json.dumps(member, 
sort_keys=True))
+    if isinstance(value, Mapping):
+        if not isinstance(rendered, dict):
+            return rendered
+        if len(rendered) != len(value):
+            # Distinct keys that render alike, such as 1 and "1": the digest 
could no
+            # longer tell those payloads apart, so refuse rather than hash 
either one.
+            raise TypeError("dict keys collide once rendered as JSON")
+        return {
+            key: _order_sets(item, rendered_item)
+            for (_, item), (key, rendered_item) in zip(value.items(), 
rendered.items())
+        }
+    if isinstance(value, (list, tuple)):
+        if not isinstance(rendered, list) or len(rendered) != len(value):
+            return rendered
+        return [_order_sets(item, rendered_item) for item, rendered_item in 
zip(value, rendered)]
+    if isinstance(rendered, dict) and (isinstance(value, BaseModel) or 
_is_dataclass_instance(value)):
+        return {
+            key: _order_sets(getattr(value, key), rendered_item) if 
hasattr(value, key) else rendered_item

Review Comment:
   Matching rendered keys back to attributes by name misses two shapes, and 
their sets keep hash-seed order. An aliased field renders under its alias, so 
`hasattr` fails. That covers `tags: set[str] = Field(alias="labels")` in a tool 
argument, and a camelCase model (`alias_generator=to_camel, 
serialize_by_alias=True`) returned into `ToolReturnPart.content`, where every 
later model step then misses on retry with the "changed" reason and no "could 
not fingerprint" warning. A `RootModel[set[str]]` renders straight to a list 
and never reaches this branch. Three seeds gave three digests for each. Mapping 
the rendered key back through `{f.serialization_alias or f.alias or name: name 
for name, f in type(value).__pydantic_fields__.items()}` (plain dataclasses 
fall back to the name) and recursing into `value.root` for a `RootModel` made 
all three agree across seeds when I tried it.



##########
providers/common/ai/src/airflow/providers/common/ai/durable/fingerprint.py:
##########
@@ -98,19 +123,107 @@ def _strip_volatile(messages_dump: list[dict[str, Any]]) 
-> list[dict[str, Any]]
     return stripped
 
 
+def _is_dataclass_instance(value: Any) -> bool:
+    return dataclasses.is_dataclass(value) and not isinstance(value, type)
+
+
+def _refuse_iterators(value: Any) -> None:
+    """
+    Raise ``TypeError`` if ``value`` holds an iterator anywhere inside it.
+
+    Rendering an iterator consumes it, and pydantic validates an 
``Iterable[T]``
+    tool parameter lazily into a ``ValidatorIterator``. Tool arguments are
+    fingerprinted before the tool runs, so rendering them would drain the 
tool's
+    own input: the tool would see an empty sequence, and that wrong result 
would be
+    cached under the fingerprint of the full one. Refusing degrades the step 
to the
+    not-cached path instead.
+    """
+    if value is None or isinstance(value, (str, bytes, bytearray, bool, int, 
float)):
+        return
+    if isinstance(value, Iterator):
+        raise TypeError(f"cannot fingerprint {type(value).__name__} without 
consuming it")
+    if isinstance(value, Mapping):
+        children: Iterable[Any] = value.values()
+    elif isinstance(value, (list, tuple, set, frozenset)):
+        children = value
+    elif isinstance(value, BaseModel):
+        children = [getattr(value, name, None) for name in 
type(value).model_fields]
+    elif _is_dataclass_instance(value):
+        children = [getattr(value, field.name, None) for field in 
dataclasses.fields(value)]
+    else:
+        return
+    for child in children:
+        _refuse_iterators(child)
+
+
+def _order_sets(value: Any, rendered: Any) -> Any:
+    """
+    Return ``rendered`` with every list that pydantic rendered from a set 
sorted.
+
+    ``rendered`` is pydantic's JSON rendering of ``value`` and is otherwise 
kept as
+    is, so pydantic stays the only renderer: bytes, dates, dict keys, ``NaN`` 
and
+    custom serializers come out exactly as a json-mode dump renders them. The 
one
+    thing pydantic cannot do stably is order a set. It lists the members in
+    iteration order, which for strings follows the interpreter's hash seed, and
+    every task attempt is a fresh process, so a ``set[str]`` would hash 
differently
+    on each attempt and never replay. ``value`` is walked alongside 
``rendered``
+    only to find those lists; a branch whose rendering does not line up with 
the
+    object, such as one with a custom serializer, is left exactly as rendered.
+    """
+    if isinstance(value, (set, frozenset)):
+        if not isinstance(rendered, list) or len(rendered) != len(value):
+            return rendered
+        members = [_order_sets(member, item) for member, item in zip(value, 
rendered)]
+        return sorted(members, key=lambda member: json.dumps(member, 
sort_keys=True))
+    if isinstance(value, Mapping):
+        if not isinstance(rendered, dict):
+            return rendered
+        if len(rendered) != len(value):
+            # Distinct keys that render alike, such as 1 and "1": the digest 
could no
+            # longer tell those payloads apart, so refuse rather than hash 
either one.
+            raise TypeError("dict keys collide once rendered as JSON")

Review Comment:
   A length mismatch here is not always a collision: a field serializer that 
filters a dict renders fewer keys than the object holds. A tool that returns a 
model whose `@field_serializer("headers")` drops `authorization` lands in 
`ToolReturnPart.content`, this raises, and every later model step fingerprints 
`None` with a warning about colliding keys that don't collide. Main hashed that 
history, and the rendered dict is also what the model receives. The docstring 
above says a branch like this is left as rendered. Raising only when the keys 
really collide, `len(to_jsonable_python({k: None for k in value}, 
bytes_mode="base64")) < len(value)`, and otherwise returning `rendered` keeps 
the refusal for `{1: "a", "1": "b"}` and gives main's digest back for the 
serializer case.



##########
providers/common/ai/src/airflow/providers/common/ai/durable/fingerprint.py:
##########
@@ -33,24 +33,43 @@
 
 Fields that pydantic-ai regenerates on every attempt (message-level
 ``timestamp``/``run_id``/``conversation_id`` and part-level ``timestamp``)
-are excluded from the fingerprint.  Requests that cannot be serialized to
-JSON fingerprint as ``None``, which degrades that step to unverified
-positional replay (the pre-fingerprint behavior) rather than disabling
-caching.
+are excluded from the fingerprint.  Every payload is rendered by pydantic in
+JSON mode before hashing, so values that are not JSON types but render the same
+way on every attempt -- a ``datetime`` or ``Decimal`` tool argument, a 
dataclass
+in ``tool_choice``, bytes in a ``BinaryContent`` -- still produce a usable
+fingerprint, and the message history hashes exactly as it always has.  The one
+correction applied on top is that set members are ordered by their JSON 
encoding,
+because a set of strings iterates in an order that follows the interpreter's 
hash
+seed and every task attempt is a fresh process (see ``_order_sets``).
+
+A request that cannot be rendered fingerprints as ``None``: that step is
+neither replayed nor cached, and re-runs live instead of replaying without
+verification.  A lazily validated ``Iterable`` argument lands there 
deliberately,
+since hashing it would consume the input the tool has not read yet, as does a
+value pydantic cannot serialize at all.  On the model path this is seldom
+confined to one step, because model settings, the tool definitions and the
+message history are carried into every later request, so durable execution 
stops
+contributing anything for the rest of the run.  A tool call is fingerprinted 
from its name, arguments and
+call id alone, so it can only lose its own step.

Review Comment:
   "It can only lose its own step" is the claim the rst paragraph moved away 
from last round. If the live re-run returns something different, the result is 
in the history the following model steps fingerprint, so they miss too. 
Something like "so it cannot make another step unfingerprintable, though a 
re-run that returns something different still invalidates the model steps after 
it" would match the docs.



##########
providers/common/ai/tests/unit/common/ai/durable/test_fingerprint.py:
##########
@@ -210,3 +237,339 @@ def test_arg_order_does_not_matter(self):
         assert fingerprint_tool_call("t", {"a": 1, "b": 2}, "id1") == 
fingerprint_tool_call(
             "t", {"b": 2, "a": 1}, "id1"
         )
+
+
+class TestPydanticNativeValues:
+    """Values that are not JSON types but hash the same on every attempt must 
still fingerprint.
+
+    Tool arguments reach ``fingerprint_tool_call`` already coerced by 
pydantic, and
+    ``tool_choice`` accepts a dataclass while genuinely affecting the 
response, so it
+    cannot be stripped as transport-only. Since a step that cannot be 
fingerprinted is
+    no longer cached at all, refusing these values would stop an ordinary 
typed tool
+    from ever being cached.
+    """
+
+    def test_datetime_tool_argument_fingerprints(self):
+        when = datetime.datetime(2026, 1, 1, tzinfo=datetime.timezone.utc)
+
+        assert fingerprint_tool_call("t", {"when": when}, "id1") is not None
+
+    def test_decimal_tool_argument_fingerprints(self):
+        assert fingerprint_tool_call("t", {"amount": Decimal("10.5")}, "id1") 
is not None
+
+    def test_datetime_tool_argument_is_stable_and_distinguishing(self):
+        early = datetime.datetime(2026, 1, 1, tzinfo=datetime.timezone.utc)
+        late = datetime.datetime(2026, 6, 1, tzinfo=datetime.timezone.utc)
+
+        assert fingerprint_tool_call("t", {"when": early}, "id1") == 
fingerprint_tool_call(
+            "t", {"when": early}, "id1"
+        )
+        assert fingerprint_tool_call("t", {"when": early}, "id1") != 
fingerprint_tool_call(
+            "t", {"when": late}, "id1"
+        )
+
+    def test_tool_choice_dataclass_fingerprints(self):
+        fp = fingerprint_model_request(
+            "m",
+            make_messages(),
+            {"tool_choice": ToolOrOutput(function_tools=["my_tool"])},
+            ModelRequestParameters(),
+        )
+
+        assert fp is not None
+
+    def test_tool_choice_dataclass_still_affects_the_fingerprint(self):
+        one = fingerprint_model_request(
+            "m",
+            make_messages(),
+            {"tool_choice": ToolOrOutput(function_tools=["a"])},
+            ModelRequestParameters(),
+        )
+        other = fingerprint_model_request(
+            "m",
+            make_messages(),
+            {"tool_choice": ToolOrOutput(function_tools=["b"])},
+            ModelRequestParameters(),
+        )
+
+        assert one is not None
+        assert one != other
+
+    def test_value_pydantic_cannot_serialize_still_returns_none(self):
+        """Normalization must not turn a genuinely unserializable value into a 
hash."""
+        assert fingerprint_tool_call("t", {"v": object()}, "id1") is None
+
+    def test_plain_payload_digest_is_unchanged_by_normalization(self):
+        """Fingerprints recorded before normalization must still match, so 
cached entries survive."""
+        payload = {"model": "m", "args": {"b": [1, True, None, "x", 2.5]}, 
"settings": None}
+        pre_normalization = hashlib.sha256(json.dumps(payload, 
sort_keys=True).encode()).hexdigest()
+
+        assert _digest(_render(payload)) == pre_normalization
+
+    def test_dict_keys_render_as_a_json_mode_dump_renders_them(self):
+        by_day = {datetime.date(2026, 1, 1): 1.5, datetime.date(2026, 1, 2): 
2.5}
+
+        assert _render({"series": by_day}) == {"series": {"2026-01-01": 1.5, 
"2026-01-02": 2.5}}
+        assert _render({1: "a", "b": 2}) == {"1": "a", "b": 2}
+
+    def test_keys_that_collide_once_rendered_are_refused(self):
+        """``1`` and ``"1"`` render alike, so hashing either payload could 
replay the other."""
+        assert fingerprint_tool_call("t", {"d": {1: "a", "1": "b"}}, "id1") is 
None
+
+    def test_bytes_that_are_not_utf8_render_as_base64(self):
+        """Decoding bytes as text would raise on an image and stop the tool 
from being cached."""
+        assert _render({"image": _PNG}) == {"image": 
base64.urlsafe_b64encode(_PNG).decode()}
+        assert fingerprint_tool_call("t", {"image": _PNG}, "id1") is not None
+
+
+def _json_mode_reference(model_identifier, messages, model_request_parameters):
+    """The fingerprint main computes: pydantic's json-mode dump, hashed as is.
+
+    For anything that is not a set, fingerprints must equal this, or entries 
stored
+    by an earlier version stop matching and the first retry after an upgrade 
re-runs
+    the whole agent.
+    """
+    dumped = ModelMessagesTypeAdapter.dump_python(messages, mode="json")
+    stripped = [
+        {
+            **{k: v for k, v in message.items() if k not in ("timestamp", 
"run_id", "conversation_id")},
+            "parts": [{k: v for k, v in part.items() if k != "timestamp"} for 
part in message["parts"]],
+        }
+        for message in dumped
+    ]
+    params = 
TypeAdapter(ModelRequestParameters).dump_python(model_request_parameters, 
mode="json")
+    payload = {"model": model_identifier, "messages": stripped, "settings": 
None, "params": params}
+    return hashlib.sha256(json.dumps(payload, 
sort_keys=True).encode()).hexdigest()
+
+
+def _with_tool_return(content):
+    return [
+        ModelRequest(parts=[UserPromptPart(content="go")]),
+        ModelResponse(parts=[ToolCallPart(tool_name="t", args={}, 
tool_call_id="c1")]),
+        ModelRequest(parts=[ToolReturnPart(tool_name="t", content=content, 
tool_call_id="c1")]),
+    ]
+
+
+class TestMessageHistoryMatchesJsonModeDump:
+    """The message history must hash exactly as pydantic's json-mode dump 
renders it.
+
+    Each case here is something a hand-written renderer gets wrong: raw bytes 
(not
+    valid UTF-8), dict keys that are not strings, ``NaN``, tuples, and fields 
whose
+    serializer only applies in JSON mode, such as ``InstructionPart.id``.
+    """
+
+    @pytest.mark.parametrize(
+        "messages",
+        [
+            pytest.param(
+                [
+                    ModelRequest(
+                        parts=[
+                            UserPromptPart(content=["look", 
BinaryContent(data=_PNG, media_type="image/png")])
+                        ]
+                    )
+                ],
+                id="binary-content-in-prompt",
+            ),
+            pytest.param(_with_tool_return(_PNG), id="bytes-tool-return"),
+            pytest.param(
+                _with_tool_return({datetime.date(2026, 1, 1): 1.5, 
datetime.date(2026, 1, 2): 2.5}),
+                id="date-keyed-tool-return",
+            ),
+            pytest.param(_with_tool_return({1: "a", "b": 2}), 
id="mixed-key-tool-return"),
+            pytest.param(_with_tool_return({"x": math.nan}), 
id="nan-tool-return"),
+            pytest.param(
+                _with_tool_return({"when": datetime.datetime(2026, 1, 1)}), 
id="datetime-tool-return"
+            ),
+        ],
+    )
+    def test_fingerprint_equals_the_json_mode_digest(self, messages):
+        fp = fingerprint_model_request("m", messages, None, 
ModelRequestParameters())
+
+        assert fp is not None
+        assert fp == _json_mode_reference("m", messages, 
ModelRequestParameters())
+
+    def test_request_parameters_hash_as_their_json_mode_dump(self):
+        """An agent's instructions reach the request parameters with a 
json-only serializer."""
+        seen = {}
+
+        class Spy(FunctionModel):
+            async def request(self, messages, model_settings, 
model_request_parameters):
+                seen.setdefault("messages", messages)
+                seen.setdefault("params", model_request_parameters)
+                return await super().request(messages, model_settings, 
model_request_parameters)
+
+        async def respond(messages, info):
+            return ModelResponse(parts=[TextPart(content="ok")])
+
+        Agent(Spy(respond), instructions="Be terse.").run_sync("hi")
+
+        fp = fingerprint_model_request("m", seen["messages"], None, 
seen["params"])
+
+        assert fp == _json_mode_reference("m", seen["messages"], 
seen["params"])
+
+    def test_tuple_parts_still_drop_part_timestamps(self):
+        """Message history passed as objects can hold its parts in a tuple."""
+        t1 = datetime.datetime(2026, 1, 1, tzinfo=datetime.timezone.utc)
+        t2 = datetime.datetime(2026, 1, 2, tzinfo=datetime.timezone.utc)
+
+        def history(timestamp):
+            return [ModelRequest(parts=(UserPromptPart(content="hi", 
timestamp=timestamp),))]
+
+        assert fingerprint_model_request(
+            "m", history(t1), None, ModelRequestParameters()
+        ) == fingerprint_model_request("m", history(t2), None, 
ModelRequestParameters())
+
+    def test_set_in_a_tool_return_hashes_as_its_ordered_members(self):
+        assert fingerprint_model_request(
+            "m", _with_tool_return({"tags": {"b", "c", "a"}}), None, 
ModelRequestParameters()
+        ) == fingerprint_model_request(
+            "m", _with_tool_return({"tags": ["a", "b", "c"]}), None, 
ModelRequestParameters()
+        )
+
+
+class TestSetMemberOrdering:
+    """Sets must hash in a fixed order rather than the interpreter's iteration 
order.
+
+    A ``set[str]`` iterates in an order derived from the process hash seed, 
and every
+    task attempt runs in a fresh process. Hashing that order would produce a 
digest
+    the next attempt never reproduces, so the step would re-run live on every 
retry
+    -- worse than declining to cache it, which at least costs nothing extra.
+    """
+
+    def test_set_hashes_as_its_ordered_members(self):
+        assert _render({"tags": {"beta", "alpha"}}) == {"tags": ["alpha", 
"beta"]}
+
+    def test_set_matches_the_equivalent_list(self):
+        assert _render({"tags": {"alpha", "beta", "gamma"}}) == 
_render({"tags": ["alpha", "beta", "gamma"]})
+
+    def test_frozenset_matches_set(self):
+        assert _render({"tags": frozenset({"a", "b"})}) == _render({"tags": 
{"b", "a"}})
+
+    def test_different_members_still_produce_different_digests(self):
+        assert _render({"tags": {"a", "b"}}) != _render({"tags": {"a", "c"}})
+
+    def test_set_nested_inside_a_list(self):
+        assert _render({"filters": [{"z", "y"}]}) == {"filters": [["y", "z"]]}
+
+    def test_set_inside_a_dataclass_field(self):
+        @dataclasses.dataclass
+        class Filter:
+            tags: set
+
+        assert _render(Filter(tags={"b", "a"})) == {"tags": ["a", "b"]}
+
+    def test_set_inside_a_basemodel_field(self):
+        class Filter(pydantic.BaseModel):
+            tags: set[str]
+
+        assert _render(Filter(tags={"b", "a"})) == {"tags": ["a", "b"]}
+
+    def test_digest_is_stable_across_process_hash_seeds(self):
+        """The real proof: two fresh interpreters must agree, as two attempts 
would.
+
+        In-process comparisons cannot catch a hash-seed dependency, since one 
process
+        has one seed. The subprocess loads the module by path so it does not 
pay for
+        importing Airflow.
+        """
+        snippet = (
+            "import importlib.util;"
+            f"spec = importlib.util.spec_from_file_location('fp', 
r'{fingerprint_module.__file__}');"
+            "mod = importlib.util.module_from_spec(spec);"
+            "spec.loader.exec_module(mod);"
+            "print(mod._digest(mod._render({'tags': {'alpha', 'beta', 'gamma', 
'delta'}})))"
+        )
+        digests = set()
+        for seed in ("0", "1", "2", "42"):
+            completed = subprocess.run(
+                [sys.executable, "-c", snippet],
+                capture_output=True,
+                text=True,
+                env={**os.environ, "PYTHONHASHSEED": seed, "PYTHONWARNINGS": 
"ignore"},
+                check=True,
+            )
+            digests.add(completed.stdout.strip().splitlines()[-1])
+
+        assert len(digests) == 1, f"digest depends on the hash seed: {digests}"
+
+    def test_set_returned_by_a_tool_is_stable_across_process_hash_seeds(self):
+        """The same proof through the message history, where a tool's set 
return lands."""
+        snippet = (
+            "import importlib.util;"
+            "from pydantic_ai.messages import ModelRequest, ToolReturnPart;"
+            "from pydantic_ai.models import ModelRequestParameters;"
+            f"spec = importlib.util.spec_from_file_location('fp', 
r'{fingerprint_module.__file__}');"
+            "mod = importlib.util.module_from_spec(spec);"
+            "spec.loader.exec_module(mod);"
+            "part = ToolReturnPart(tool_name='t', content={'alpha', 'beta', 
'gamma', 'delta'}, tool_call_id='c1');"
+            "print(mod.fingerprint_model_request('m', 
[ModelRequest(parts=[part])], None, ModelRequestParameters()))"
+        )
+        digests = set()
+        for seed in ("0", "1", "2", "42"):
+            completed = subprocess.run(
+                [sys.executable, "-c", snippet],
+                capture_output=True,
+                text=True,
+                env={**os.environ, "PYTHONHASHSEED": seed, "PYTHONWARNINGS": 
"ignore"},
+                check=True,
+            )
+            digests.add(completed.stdout.strip().splitlines()[-1])
+
+        assert "None" not in digests
+        assert len(digests) == 1, f"digest depends on the hash seed: {digests}"
+
+
+class TestLazilyValidatedIterable:
+    """A lazily validated ``Iterable`` must be refused, not consumed.
+
+    ``to_jsonable_python`` drains an iterator to render it. Tool arguments are
+    fingerprinted before the tool runs, so draining one would hand the tool an
+    exhausted iterator and cache that wrong result under the fingerprint of the
+    full input. Declining to fingerprint leaves the step uncached instead.
+    """
+
+    def test_validator_iterator_argument_is_not_fingerprinted(self):

Review Comment:
   Every case in this class passes the iterator at the top level of the args. 
Deleting the list/tuple, `BaseModel` or dataclass branch of 
`_refuse_iterators`, or the `_refuse_iterators(messages)` call, leaves the 
durable suite green, and each deletion brings back the drained-input bug for a 
nested parameter such as `list[Iterable[int]]` or a model field `ids: 
Iterable[int]`. Parametrizing over those shapes (plus the extras case) and 
adding a `fingerprint_model_request` case with a generator in 
`ToolReturnPart.content` would pin them. The generator case is also a fix worth 
recording: on main the model received `[]` there. Removing the `_order_sets` 
wrap on the request parameters is likewise not caught, even though 
`revealed_tool_names` is a `set[str]` in real requests. Putting a few names 
there in the subprocess seed test would guard it.



##########
providers/common/ai/docs/durable_execution.rst:
##########
@@ -102,6 +102,33 @@ cache:
    never replays responses that belong to a different conversation.
 4. After successful completion, the cached steps are deleted.
 
+Fingerprints are computed from pydantic's JSON rendering of each value, the 
same
+rendering a json-mode dump produces, so ordinary types that are not JSON -- a
+``datetime`` or ``Decimal`` tool argument, a dataclass in ``tool_choice``, the
+bytes in a ``BinaryContent``, a dict keyed by date -- fingerprint normally, and
+entries cached by an earlier version still match. The one adjustment is that 
the
+members of a ``set`` are ordered before hashing, so a set matches on a later
+attempt too. If a value cannot be rendered at all, that step is not cached, 
and on
+retry it runs live rather than replaying an unverified entry. A parameter 
annotated ``Iterable[...]`` is one such case:
+pydantic validates it lazily, and reading it in order to hash it would consume
+the input the tool itself has not read yet, so the step runs live instead of
+being cached.
+
+On the model path this is rarely confined to a single step: the causes are 
such a
+value in ``model_settings``, which is attached to every request, in the tool
+definitions the request carries, or in the message history, which every later
+request carries forward. Any one of them degrades all subsequent model steps 
the
+same way, leaving durable execution with nothing to
+replay, so the retry re-runs the agent at full cost. The

Review Comment:
   "Nothing to replay" and "re-runs the agent at full cost" hold when the cause 
is in `model_settings` or the tool definitions, since those are there from the 
first request. When it enters the history mid-run, from a tool return for 
example, the steps before it still replay and only the steps after it run live. 
Something like "degrades every model step from that point on, so the retry 
re-runs the agent from there at full cost, and from the start when the cause is 
in the settings or tool definitions" would cover both.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to