weiqingy opened a new pull request, #1215:
URL: https://github.com/apache/flink-agents/pull/1215

   Linked issue: #912
   
   ### Purpose of change
   
   When an agent asks for structured output and its chat model can enforce a 
schema natively, the final answer now comes from a provider call that carries 
the schema, so the provider enforces the format instead of the model following 
a prompt. This completes #280: every provider can apply a schema natively, but 
until now the agent runtime never handed one over.
   
   Nothing changes for a request whose schema does not resolve to native. Under 
the default `AUTO` strategy that covers every model a connection does not list 
as honoring native structured output, and every `RowTypeInfo` schema.
   
   #### Runtime flow
   
   1. `ReActAgent` asks the default chat model's setup whether the schema will 
travel natively. If yes, it leaves the "match the schema" instruction out of 
the prompt.
   2. The chat model action asks the same question once per request, before its 
retry loop.
   3. The tool-calling loop runs as today.
   4. When the model's answer has no tool calls and the answer above was yes, 
the same attempt makes one more call: the same request history, the answer, and 
a short user instruction to convert that answer into the required format. This 
call carries the schema and no tools. Its response is the one parsed.
   
   #### Key decisions
   
   - **The extra call happens only when the schema resolves to native**, as 
agreed on #912. The prompt path keeps today's single call.
   - **The extra call binds no tools**, because Gemini drops a native schema 
when tools are bound. It also leaves tool calls and tool results out of the 
history, which Anthropic and Bedrock reject in a request that defines no tools.
   - **The setup's strategy is the policy; a connection only encodes.** A 
connection now carries the schema whenever the request can carry it, without 
also checking its model list. Otherwise `NATIVE` on a model outside that list 
would silently lose the schema.
   
   ### Behavioral Semantics
   
   #### Interaction decisions
   
   | Strategy | Connection's answer | Result |
   |---|---|---|
   | `PROMPT` | any | Prompt path, no extra call |
   | `AUTO` | native recommended | Native: extra call, instruction omitted |
   | `AUTO` | feasible, or infeasible | Prompt path |
   | `NATIVE` | native recommended, or feasible | Native |
   | `NATIVE` | infeasible (e.g. a `RowTypeInfo` schema) | Request fails before 
any model call |
   | `AUTO` / `PROMPT` | setup has no connection, or is bridged to the other 
language | Prompt path |
   | `NATIVE` | setup has no connection, or is bridged to the other language | 
Request fails before any model call |
   
   #### Behavioral contracts
   
   1. The extra call is made only when the strategy resolves to native and the 
answer has no tool calls.
   2. Its messages are the history the loop call sent, then the answer, then a 
user message with the conversion instruction.
   3. It carries the schema, binds no tools, and contains no tool calls or tool 
results.
   4. Its response is the one parsed into structured output.
   5. A `NATIVE` strategy that cannot be honored fails the request with a 
message naming the connection (or setup) and the schema.
   6. `ReActAgent` omits the schema instruction only when the answer is native; 
any other answer, or any failure to get one, keeps it.
   7. A recovered run replays both calls from the durable record without 
calling the model.
   
   #### Failure behavior
   
   - If the extra call fails, is truncated, or does not parse, it uses the 
existing retry budget; a retry repeats the attempt, not the completed tool 
rounds.
   - A `NATIVE` strategy that cannot be honored fails like any other model 
failure: in Java a router falls back to its next candidate, otherwise a failed 
`ChatResponseEvent` is sent.
   - Cancellation propagates everywhere, including from the `ReActAgent` check.
   
   ### Tests
   
   | Contract | Pinned by |
   |---|---|
   | 1 | `ChatModelActionRetryTest`, `ChatModelInvokerTest`, 
`test_chat_model_action_retry.py` |
   | 2 | `finalizationSendsThePreparedRequestNotTheRawInput`; the real-setup 
history test in `test_chat_model_action_retry.py` |
   | 3 | `BaseChatModelTest` tool-traffic tests; `test_chat_model_base.py` 
twins |
   | 4 | parse-from-final tests in both languages |
   | 5 | strategy x answer grids in `BaseChatModelTest` and 
`test_chat_model_base.py`; bridge tests |
   | 6 | `ReActAgentTest`; `test_react_agent_schema_instruction.py` |
   | 7 | `ChatModelInvokerTest` replay; 
`test_flink_runner_context_reconcilable.py` (sync and async, across a retry) |
   
   Each native connection also tests that a model outside its list still 
carries the schema. Every commit was mutation-tested, and all non-equivalent 
mutants were killed.
   
   **Not verified:**
   - No call to a live provider. The Anthropic and Bedrock rejections of tool 
traffic without tools are taken from their API documentation.
   - The setup's answer is not stored in the durable record. If configuration 
changes between a crash and recovery, the replay may take the other path; the 
journal then discards the later records, so the result is a fresh valid answer.
   
   ### API
   
   - New on `BaseChatModelSetup`: `willApplyNativeStructuredOutput` / 
`chatStructured` (Java) and `will_apply_native_structured_output` / 
`chat_structured` / `prepare_request_messages` (Python). Users configure 
behavior through `structured_output_strategy`, not by calling these.
   - **Behavior change for direct callers of `connection.chat(..., 
outputSchema)`:** a model outside a connection's native list now gets the 
native parameter. On that path Anthropic drops its JSON prefill, Azure and 
DashScope reject a caller-supplied `response_format`, and Python raises 
`TypeError` for a schema Pydantic cannot render instead of falling back to the 
prompt.
   - The connection contract from #1129 is reworded accordingly: the native 
branch carries the schema whenever the connection's answer is not infeasible.
   - Agent plans serialize unchanged.
   
   ### Documentation
   
   - [x] `doc-needed`
   - [ ] `doc-not-needed`
   - [ ] `doc-included`
   
   User documentation for structured output across providers follows in a 
separate PR.
   
   ### Was this patch authored or co-authored using generative AI tooling?
   
   - [x] Yes
   - [ ] No
   
   Generated-by: Claude Code 2.1.294 (Claude Opus 5.5)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to