Zhuoxi2000 opened a new pull request, #1178:
URL: https://github.com/apache/flink-agents/pull/1178

   <!--
   * Thank you very much for contributing to Flink Agents.
   * Please add the relevant components in the PR title. E.g., [api], 
[runtime], [java], [python], [hotfix], etc.
   -->
   
   <!-- Please link the PR to the relevant issue(s). Hotfix doesn't need this. 
-->
   Linked issue: #1059
   
   ### Purpose of change
   
   <!-- What is the purpose of this change? -->
   
   A user `ChatMessage` with image blocks now reaches Ollama with its images, 
in Java and Python, so vision models such as `qwen2.5vl` or `llava` can see 
them. Until now the Ollama connections sent only the text and dropped the 
images silently. This is the second provider step of #1059 Phase 2, after #1164.
   
   #### Runtime flow
   
   1. Each message is converted as before (`convertToOllamaChatMessages` in 
Java, `__convert_to_ollama_messages` in Python), including assistant tool calls 
(forwarded in Java since #1166).
   2. The message's media blocks are then checked in order: a user message's 
Base64 `ImageBlock`s are decoded and attached as the message's `images`; any 
other media throws before the request is sent.
   3. A message without media is sent exactly as before, with no `images` field.
   
   #### Key decisions
   
   * Images are attached next to the text, not interleaved: Ollama's 
`/api/chat` has one `content` string and one `images` list per message, so 
block order is kept among the images but not between text and images.
   * Only Base64 images: Ollama takes inline image data and cannot fetch a URL, 
and fetching it in the connection would make the framework an HTTP client for 
arbitrary URLs.
   * The Base64 data is decoded to bytes. ollama4j takes bytes and re-encodes 
them; ollama-python treats a string as a file path first, so bytes avoid an 
accidental file read.
   * Unsupported media raises the shared `UnsupportedContentBlockException` / 
`UnsupportedContentBlockError` from #1164.
   
   ### Behavioral Semantics
   
   <!-- For a non-trivial code change whose implementation is largely 
AI-assisted: interaction decisions, behavioral contracts, and failure behavior. 
See `contribution-guides/ai-assisted-pr.md`. Remove this heading and this 
comment otherwise. -->
   
   #### Interaction decisions
   
   | Role | Blocks | Result |
   |---|---|---|
   | user | text only, or none | sent as before, no `images` |
   | user | Base64 images, with or without text | `content` = text projection, 
`images` = the images in block order |
   | user | an image by URL, audio, video or a document | 
`UnsupportedContentBlockException`; nothing sent |
   | system / assistant / tool | any media | 
`UnsupportedContentBlockException`; nothing sent |
   
   #### Behavioral contracts
   
   1. A message without media is sent unchanged, without `images`.
   2. A user message's Base64 images are sent as `images`, in block order, and 
serialize back to the original Base64 strings.
   3. The message `content` is the text projection.
   4. An image by URL, and any audio, video or document block, throws 
`UnsupportedContentBlockException` naming the block but not its data or URL.
   5. Media in a system, assistant or tool message throws 
`UnsupportedContentBlockException`.
   6. Image data that is not valid Base64 throws `IllegalArgumentException` / 
`ValueError` ("could not be decoded"), not the unsupported-block error.
   
   Java and Python behave identically for each contract.
   
   #### Failure behavior
   
   * Unsupported media or role, and undecodable image data: thrown while 
building the request; nothing is sent, and the chat action's error strategy 
applies. These errors are deterministic, so RETRY fails the same way each 
attempt.
   * A model without vision support: Ollama's own error propagates unchanged.
   
   ### Tests
   
   <!-- How is this change verified? -->
   
   | Contract | Java: `OllamaMultimodalTest` | Python: 
`test_ollama_multimodal.py` |
   |---|---|---|
   | 1 | `testTextOnlyMessageHasNoImages` | 
`test_text_only_message_has_no_images` |
   | 2, 3 | `testBase64ImagesAttachedInBlockOrder` | 
`test_base64_images_attached_in_block_order` |
   | 4 | `testUnsupportedMediaFailsExplicitly` | 
`test_unsupported_media_fails_explicitly` |
   | 5 | `testImagesOutsideUserMessagesFail` | 
`test_images_outside_user_messages_fail` |
   | 6 | `testInvalidBase64Fails` | `test_invalid_base64_fails` |
   
   Java asserts on the request as ollama4j serializes it; Python asserts on the 
messages passed to the mocked client, serialized. Opt-in live tests 
(`OllamaMultimodalLiveTest`, `test_ollama_multimodal_live.py`) send a real 
image when `OLLAMA_VISION_MODEL` is set.
   
   Not verified:
   
   * A live Ollama server: no vision model was available here, so the live 
tests were not run.
   * Interleaving text and images: not representable in Ollama's API (see Key 
decisions).
   
   <details>
   <summary>Implementation invariants and supporting evidence</summary>
   
   * ollama4j 1.1.5 `OllamaChatMessage.images` is `List<byte[]>` serialized 
with `FileToBase64Serializer`; ollama-python 0.6.1 `Message.images` holds 
`Image` values whose serializer base64-encodes bytes and treats strings as 
paths first.
   * The Java conversion keeps #1166's tool-call forwarding for assistant 
messages alongside the images.
   * Verified locally: Ollama module tests 23/23 with spotless; Python Ollama 
and chat-message tests pass, with ruff.
   
   </details>
   
   ### API
   
   <!-- Does this change touches any public APIs? -->
   
   No new API; it uses the `UnsupportedContentBlockException` / 
`UnsupportedContentBlockError` added in #1164.
   
   Compatibility: text-only requests are unchanged. User images used to be 
dropped silently; they are now sent. Other media, and media in non-user 
messages, now fail instead of being dropped.
   
   ### Documentation
   
   <!-- Do not remove this section. Check the proper box only. -->
   
   - [ ] `doc-needed` <!-- Your PR changes impact docs -->
   - [ ] `doc-not-needed` <!-- Your PR changes do not impact docs -->
   - [x] `doc-included` <!-- Your PR already contains the necessary 
documentation updates -->
   
   A multimodal note in the Ollama section of `chat_models.md`, and Ollama 
added to the provider-support note.
   
   ### Was this patch authored or co-authored using generative AI tooling?
   
   <!-- Do not remove this section. Check the proper box only. -->
   
   - [x] Yes
   - [ ] No
   
   If yes, include a `Generated-by: <tool name and version> (<model name and 
version>)` line, for example `Generated-by: Claude Code 2.1.226 (Claude Opus 
4.6)`, in the commit message so it reaches Git history. Repeat the same line 
here for reviewer visibility. See the [ASF generative tooling 
guidance](https://www.apache.org/legal/generative-tooling.html).
   
   Generated-by: Claude Code 2.1.259 (Claude Opus 5.5)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to