weiqingy opened a new issue, #912:
URL: https://github.com/apache/flink-agents/issues/912

   ### Search before asking
   
   - [X] I searched in the 
[issues](https://github.com/apache/flink-agents/issues) and found nothing 
similar.
   
   ### Description
   
   Follow-up to #280. This issue covers applying structured output to an 
agent's **final output only**, which #280 scoped as a separate step from the 
chat-model-layer foundation being added in #843.
   
   **Problem.** When a `ReActAgent` declares an `output_schema`, the schema is 
honored purely by prompt engineering today: a system prompt instructs the model 
to emit JSON, and the response text is fence-stripped and parsed afterwards. At 
the chat-model layer, an intermediate tool-calling turn and the final answering 
turn are indistinguishable, so there is no point at which structured output can 
be applied to just the final result. The #843 foundation adds an explicit 
output-schema parameter to the chat path plus a policy 
(`auto`/`native`/`prompt`) and connection-side capability split, but no 
framework path passes a schema on that call yet — supplying it at the right 
moment is what this issue adds.
   
   **Proposed solution.** Issue a dedicated structured call once the loop has 
produced its final answer — at `ChatModelAction`'s no-tool-calls branch, the 
point where the model stops requesting tools. That call passes the output 
schema explicitly, so native structured output applies to the final output only 
and intermediate tool-calling turns are untouched.
   
   **Schema instructions must become strategy-driven.** `ReActAgent` currently 
registers a schema-instruction prompt (`The final response should be json 
format, and match the schema ...`) unconditionally at construction whenever 
`output_schema` is set, in both languages:
   
   - Java — `ReActAgent.java:66-87` (registered at `:86`)
   - Python — `react_agent.py:128-139`
   
   That prompt is then injected into the messages when the action runs (Java 
`ReActAgent.java:156-162`, Python `react_agent.py:189-196`).
   
   If the finalization call uses a model that applies structured output 
natively, the prompt instruction and the provider's own schema enforcement 
become two overlapping channels describing the same schema. The 
schema-instruction path therefore needs to follow the selected strategy: emit 
it for the prompt-based fallback, omit it when the native path applies. Both 
sites are candidates for carrying that decision — the injection is per-call 
while the registration happens at construction, before the effective model is 
known — so which one carries it is part of this issue's design rather than a 
settled choice.
   
   **Open questions.**
   
   - Should finalization be an extra post-loop model round-trip, or a forced 
tool call that avoids the additional call? An extra call is the simpler 
starting point; the forced-tool form is an optimization that could follow.
   - Where should `auto` resolve policy into a concrete strategy, given 
capability is a connection-side predicate over the effective model at 
request-build time?
   
   **Scope.** Depends on the #843 foundation landing first, since it calls that 
explicit-schema chat path. Java and Python move together.
   
   ### Are you willing to submit a PR?
   
   - [X] I'm willing to submit a PR!
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to