kaxil commented on issue #73658: URL: https://github.com/apache/airflow/issues/73658#issuecomment-5879009927
Verified common.ai 0.10.0rc1 from the PyPI wheel on Airflow 3.3.2 (Python 3.12), plus Airflow 3.0.6 for the version gates, with live runs against Claude, TypeSafe Jev and Modal sandboxes. Everything below works as described in the PRs: - #73368 decision policy: Jev confidence gates `LLMBranchOperator`/`LLMOperator`; `on_uncertain="fail"` raises `LowConfidenceError`; `"review"` opens the approval flow, and accept, modify, reject and timeout each finalize the decision record correctly. A model with no confidence counts as under the bar. - #73363 TypeSafe Jev: `typesafe:` models resolve through `PydanticAIHook`; both classifier example Dags run live. - #73501 `ClassifierRetryPolicy` + `fallback_policy`: the classifier decides, escalates to `LLMRetryPolicy` below the bar or on a model error, and falls back to rules. `LLMRetryPolicy`'s signature and `ErrorClassification` schema are identical to 0.9.0. Task-level `retry_policy` retries after the category delay in a real Dag run. - #72910 Modal sandbox backend: command, file I/O, timeout and teardown work; a Claude agent wrote and ran a script in the sandbox. - #73534 CIDR egress allowlist: an allowlisted address is reachable and others time out; the refused combinations and input validation behave as documented; `sbx` refuses a CIDR list. - #73529: `durable=True` / HITL review with a `SandboxToolset` fails at parse, including through `.prefixed()` and `CombinedToolset`; a bad region, image or Modal credentials fail the task without a model retry. - #73261 / #73052: on 3.0.6, approval and review arguments fail at parse with the "needs Airflow 3.1+" message first; plain `LLMOperator` still runs on 3.0.6. - #73511: plain and structured output work on anthropic 1.8.0 and 1.7.0. - #71403: cost limit trips with `UsageLimitExceeded`, stops the agent after its first request, and resets per run and per retry. - #72938: batch mode runs end to end with only the SDK batch layer stubbed (deferral, trigger, results, re-attach on retry without resubmitting). No live provider batch run. - #72157 / #72159 / #72155: assigned reviewers (non-reviewer gets 403), notifiers (a failing one leaves the review open), and `on_approval_timeout` approve/reject/fail. One non-blocking follow-up, which I'll open a PR for: #73273 updated the `LLMSchemaCompareOperator` docs and example to use a `DataSourceConfig` with only `conn_id` + `table_name`, which needs common.sql 2.2.0, while the `sql` extra still allows common.sql 2.1.1. #73273 and #73370 themselves ship in common.sql 2.2.0rc1, so I didn't test them here. +1 for common.ai 0.10.0rc1. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
