kaxil opened a new pull request, #73368:
URL: https://github.com/apache/airflow/pull/73368

   A classifier model such as TypeSafe's Jev returns a confidence with every 
answer, and until now no operator read it. `LLMBranchOperator` branched on a 
0.52/0.48 split exactly as it branched on 0.99, and the number was dropped on 
the floor. The one workflow a classifier model exists for, "route automatically 
when sure, ask a person otherwise", was reachable only by calling the hook from 
a `@task` and reading `provider_details` by hand.
   
   This adds `review_below` to `LLMOperator` and `LLMBranchOperator`: when the 
model's confidence in its answer is under the bar, the answer goes to the 
human-review flow that `require_approval` already uses, with the confidence, 
the bar and the full probability distribution in the review form. A number is 
one bar for everything; a mapping is a bar per branch (or per output field), so 
a branch whose wrong pick costs more can demand more certainty.
   
   ```python
   LLMBranchOperator(
       task_id="triage_failure",
       prompt=traceback,
       llm_conn_id="jev_default",
       model_id="typesafe:jev-1.13.0",
       branch_descriptions={...},
       review_below={"page_oncall": 0.9, "rerun": 0.6, "ignore": 0.6},
       allow_modifications=True,
   )
   ```
   
   Both operators also push a `decision` XCom next to the result: the model 
that answered, what it proposed, what ran, the per-field confidence and 
probabilities, the bar that applied, why the answer went to review, and who 
decided. It is written before a review opens and overwritten when the review 
resolves, so a rejected or reviewer-changed pick is recorded as such.
   
   Stacked on #73367 (branch descriptions) and #73366 (option order); this 
change builds on the same lines.
   
   ## Design rationale
   
   **`require_approval` is untouched and always wins.** It means "every answer 
goes to a person". `review_below` is the conditional setting and the two are 
never folded together: `require_approval=True` with `review_below=0.5` still 
sends a 0.99 pick to review, and the record says `"review": "require_approval"`.
   
   **What happens when there is no confidence to compare is configured, and 
defaults to asking.** A text model reports no confidence, and so does a bounded 
`float` field, where the probability is the answer. With no bar configured 
nothing changes. With a bar configured and no confidence reported, 
`on_missing_confidence` decides, and the default is `"review"`: swapping the 
connection to a model that reports nothing must not silently switch off a 
control the author set. `"fail"` and `"proceed"` cover authors who want the 
other two behaviours.
   
   **Confidence is described as what it is.** pydantic-ai's TypeSafe adapter 
reports it from the shape of the probability distribution: concentrated on one 
option is high, spread out is low. It is not the probability that the answer is 
correct, and neither the docs nor the review form say it is. The full 
distribution goes into the record so a reader can compute a different statistic.
   
   **The bar is compared once, in one place.** `utils/decision.py` holds the 
reading of `provider_details`, the threshold selection (a mapping picks the 
strictest bar among what was picked when `allow_multiple_branches=True`, and a 
picked branch missing from the mapping has no bar), and the review reason. The 
operators call it and act. On `LLMOperator` a structured output is gated per 
field, using the least confident field against its bar.
   
   **The gate is not model fallback.** It runs after the model has answered and 
decides only whether the operator acts. Handing a low-confidence request to a 
second model belongs inside the model request, through pydantic-ai's 
`FallbackModel` response handlers; a gate after `run_sync()` would be too late 
to protect a tool that already ran, and re-running would repeat side effects. 
Wiring `FallbackModel` from Airflow connections is a separate change.
   
   **The record is finalised before `skip()` on rejection.** On the Task SDK, 
`skip()` hands the skip to the supervisor by raising, so nothing after it runs. 
The first live run of the reject path showed the record left as pending; the 
fix and a test in which `skip` raises are in this change.
   
   _Evidence from the live runs and screenshots follow in an edit once the PR 
exists; the body is over the browser URL limit._
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to