Zhuoxi2000 opened a new pull request, #1164: URL: https://github.com/apache/flink-agents/pull/1164
<!-- * Thank you very much for contributing to Flink Agents. * Please add the relevant components in the PR title. E.g., [api], [runtime], [java], [python], [hotfix], etc. --> <!-- Please link the PR to the relevant issue(s). Hotfix doesn't need this. --> Linked issue: #1059 ### Purpose of change <!-- What is the purpose of this change? --> A user `ChatMessage` with image, audio or document blocks now reaches OpenAI, Azure OpenAI and vLLM with its media, in Java and Python. Until now every provider sent only the text, silently dropping media. This is the first provider step of #1059 Phase 2; Ollama and explicit errors for the remaining providers follow in separate PRs. #### Runtime flow 1. The three connections share one Chat Completions converter per language: `OpenAIChatCompletionsUtils.convertToOpenAIMessage` (Java) and `convert_to_openai_message` (Python). 2. A system, assistant or tool message is first checked for media blocks and rejected if it has any. 3. A user message without media is sent with its text projection as a string, as before. With media, every block is mapped in order to a content part, and the first block without a part throws. #### Key decisions * Text-only user messages keep string content, so existing requests and OpenAI-compatible servers see no change. * Unsupported blocks throw a new `UnsupportedContentBlockException` / `UnsupportedContentBlockError` (an `IllegalArgumentException` / `ValueError`) from the api module, rather than being dropped or converted; the later provider PRs reuse it. Its message names the block type, media type and source type, never the payload or URL. * Base64 images and documents are sent as `data:` URIs, the form OpenAI documents for `image_url` and `file_data`. * Documents are not restricted to PDF: OpenAI rejects other types itself; compatible servers may accept more. * Video, audio by URL and documents by URL are rejected: Chat Completions has no part for them. vLLM's `video_url` extension is left out. ### Behavioral Semantics <!-- For a non-trivial code change whose implementation is largely AI-assisted: interaction decisions, behavioral contracts, and failure behavior. See `contribution-guides/ai-assisted-pr.md`. Remove this heading and this comment otherwise. --> #### Interaction decisions | Role | Blocks | Result | |---|---|---| | user | text only, or none | `content` is the text projection string | | user | media, every block mappable | `content` is a part list in block order, text blocks included | | user | a block without a part | `UnsupportedContentBlockException`; no request is sent | | system / assistant / tool | text only | unchanged | | system / assistant / tool | any media | `UnsupportedContentBlockException`; no request is sent | #### Behavioral contracts 1. A user message without media is sent with string content equal to its text projection. 2. A user message with media is sent as content parts in block order, one `text` part per `TextBlock`. 3. `ImageBlock` becomes `image_url` with the URL, or `data:<media_type>;base64,<data>`. 4. `AudioBlock` with Base64 data becomes `input_audio`, format `wav` for `audio/wav` (`audio/wave`, `audio/x-wav`, `audio/vnd.wave`) and `mp3` for `audio/mpeg` (`audio/mp3`); media-type parameters and case are ignored. 5. `DocumentBlock` with Base64 data becomes `file` with `file_data` as a data URI and `filename` set to the block's `name`, else `document`. 6. `VideoBlock`, audio or documents by URL, and other audio types throw `UnsupportedContentBlockException`. 7. Media in a system, assistant or tool message throws `UnsupportedContentBlockException`. 8. The exception message never contains the Base64 data or the URL. Java and Python behave identically for each contract. #### Failure behavior * Unsupported block or role: thrown while building the request, so nothing is sent; the chat action's error strategy applies as for any provider error. The error is deterministic, so RETRY fails the same way on every attempt. * A model or server that rejects an accepted part (a non-vision model, a non-PDF document on OpenAI): the provider's error propagates unchanged. * A tool message with media fails on the media before the existing `externalId` check. ### Tests <!-- How is this change verified? --> | Contract | Java: `OpenAIChatCompletionsMultimodalTest` | Python: `test_openai_multimodal.py` | |---|---|---| | 1 | `testTextOnlyUserMessageKeepsStringContent` | `test_text_only_user_message_keeps_string_content` | | 2, 3 | `testUserMediaBecomesOrderedContentParts` | `test_user_media_becomes_ordered_content_parts` | | 4 | `testAudioBecomesInputAudio` | `test_audio_becomes_input_audio` | | 5 | `testDocumentBecomesFilePart` | `test_document_becomes_file_part` | | 6, 8 | `testUnsupportedUserBlocksFailExplicitly` | `test_unsupported_user_blocks_fail_explicitly` | | 7 | `testMediaOutsideUserMessagesFails` | `test_media_outside_user_messages_fails` | Java tests assert on the request as serialized by the SDK's own mapper, i.e. the wire JSON; Python tests assert on the request dicts. Not verified: * A live OpenAI, Azure OpenAI or vLLM endpoint: no multimodal request was sent to a real model. * The audio aliases `audio/wave`, `audio/x-wav`, `audio/vnd.wave` and `audio/mp3`, and media-type parameters: mapped in code, not individually tested. * `image_url.detail` is never set, so the provider default applies. <details> <summary>Implementation invariants and supporting evidence</summary> * The converter is shared: Java `OpenAICompletionsConnection`, `AzureOpenAIChatModelConnection` and `VLLMChatModelConnection` (a subclass of the first); Python `OpenAIChatModelConnection`, `AzureOpenAIChatModelConnection` and `VLLMChatModelConnection` likewise. * openai-java 4.8.0 takes media parts on user messages only (`contentOfArrayOfContentParts`); the system, tool and assistant builders accept text parts only, which is why media is rejected by role. * `input_audio` formats in openai-java 4.8.0 and openai-python are exactly `wav` and `mp3`. * OpenAI's file-input guide: Chat Completions accepts `file_data` for PDF only, as `data:application/pdf;base64,...`. * Verified locally: the Java OpenAI integration module 132/132 and the api `ChatMessage*` tests, spotless; Python OpenAI, Azure and vLLM plus chat-message tests 191 passed (4 skipped), ruff check and format. </details> ### API <!-- Does this change touches any public APIs? --> New public types: `org.apache.flink.agents.api.chat.messages.UnsupportedContentBlockException` (Java) and `flink_agents.api.chat_message.UnsupportedContentBlockError` (Python). Compatibility: requests for text-only messages are unchanged. A user message with media used to be sent as its text only; it now carries the media, or fails if a block has no Chat Completions part. Media in a non-user message used to be dropped; it now fails. Other providers are unchanged. ### Documentation <!-- Do not remove this section. Check the proper box only. --> - [ ] `doc-needed` <!-- Your PR changes impact docs --> - [ ] `doc-not-needed` <!-- Your PR changes do not impact docs --> - [x] `doc-included` <!-- Your PR already contains the necessary documentation updates --> "Multimodal Input" under OpenAI in `chat_models.md`, linked from Azure OpenAI and vLLM. ### Was this patch authored or co-authored using generative AI tooling? <!-- Do not remove this section. Check the proper box only. --> - [x] Yes - [ ] No If yes, include a `Generated-by: <tool name and version> (<model name and version>)` line, for example `Generated-by: Claude Code 2.1.226 (Claude Opus 4.6)`, in the commit message so it reaches Git history. Repeat the same line here for reviewer visibility. See the [ASF generative tooling guidance](https://www.apache.org/legal/generative-tooling.html). Generated-by: Claude Code 2.1.259 (Claude Opus 5.5) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
