Zhuoxi2000 opened a new pull request, #1178:
URL: https://github.com/apache/flink-agents/pull/1178
<!--
* Thank you very much for contributing to Flink Agents.
* Please add the relevant components in the PR title. E.g., [api],
[runtime], [java], [python], [hotfix], etc.
-->
<!-- Please link the PR to the relevant issue(s). Hotfix doesn't need this.
-->
Linked issue: #1059
### Purpose of change
<!-- What is the purpose of this change? -->
A user `ChatMessage` with image blocks now reaches Ollama with its images,
in Java and Python, so vision models such as `qwen2.5vl` or `llava` can see
them. Until now the Ollama connections sent only the text and dropped the
images silently. This is the second provider step of #1059 Phase 2, after #1164.
#### Runtime flow
1. Each message is converted as before (`convertToOllamaChatMessages` in
Java, `__convert_to_ollama_messages` in Python), including assistant tool calls
(forwarded in Java since #1166).
2. The message's media blocks are then checked in order: a user message's
Base64 `ImageBlock`s are decoded and attached as the message's `images`; any
other media throws before the request is sent.
3. A message without media is sent exactly as before, with no `images` field.
#### Key decisions
* Images are attached next to the text, not interleaved: Ollama's
`/api/chat` has one `content` string and one `images` list per message, so
block order is kept among the images but not between text and images.
* Only Base64 images: Ollama takes inline image data and cannot fetch a URL,
and fetching it in the connection would make the framework an HTTP client for
arbitrary URLs.
* The Base64 data is decoded to bytes. ollama4j takes bytes and re-encodes
them; ollama-python treats a string as a file path first, so bytes avoid an
accidental file read.
* Unsupported media raises the shared `UnsupportedContentBlockException` /
`UnsupportedContentBlockError` from #1164.
### Behavioral Semantics
<!-- For a non-trivial code change whose implementation is largely
AI-assisted: interaction decisions, behavioral contracts, and failure behavior.
See `contribution-guides/ai-assisted-pr.md`. Remove this heading and this
comment otherwise. -->
#### Interaction decisions
| Role | Blocks | Result |
|---|---|---|
| user | text only, or none | sent as before, no `images` |
| user | Base64 images, with or without text | `content` = text projection,
`images` = the images in block order |
| user | an image by URL, audio, video or a document |
`UnsupportedContentBlockException`; nothing sent |
| system / assistant / tool | any media |
`UnsupportedContentBlockException`; nothing sent |
#### Behavioral contracts
1. A message without media is sent unchanged, without `images`.
2. A user message's Base64 images are sent as `images`, in block order, and
serialize back to the original Base64 strings.
3. The message `content` is the text projection.
4. An image by URL, and any audio, video or document block, throws
`UnsupportedContentBlockException` naming the block but not its data or URL.
5. Media in a system, assistant or tool message throws
`UnsupportedContentBlockException`.
6. Image data that is not valid Base64 throws `IllegalArgumentException` /
`ValueError` ("could not be decoded"), not the unsupported-block error.
Java and Python behave identically for each contract.
#### Failure behavior
* Unsupported media or role, and undecodable image data: thrown while
building the request; nothing is sent, and the chat action's error strategy
applies. These errors are deterministic, so RETRY fails the same way each
attempt.
* A model without vision support: Ollama's own error propagates unchanged.
### Tests
<!-- How is this change verified? -->
| Contract | Java: `OllamaMultimodalTest` | Python:
`test_ollama_multimodal.py` |
|---|---|---|
| 1 | `testTextOnlyMessageHasNoImages` |
`test_text_only_message_has_no_images` |
| 2, 3 | `testBase64ImagesAttachedInBlockOrder` |
`test_base64_images_attached_in_block_order` |
| 4 | `testUnsupportedMediaFailsExplicitly` |
`test_unsupported_media_fails_explicitly` |
| 5 | `testImagesOutsideUserMessagesFail` |
`test_images_outside_user_messages_fail` |
| 6 | `testInvalidBase64Fails` | `test_invalid_base64_fails` |
Java asserts on the request as ollama4j serializes it; Python asserts on the
messages passed to the mocked client, serialized. Opt-in live tests
(`OllamaMultimodalLiveTest`, `test_ollama_multimodal_live.py`) send a real
image when `OLLAMA_VISION_MODEL` is set.
Not verified:
* A live Ollama server: no vision model was available here, so the live
tests were not run.
* Interleaving text and images: not representable in Ollama's API (see Key
decisions).
<details>
<summary>Implementation invariants and supporting evidence</summary>
* ollama4j 1.1.5 `OllamaChatMessage.images` is `List<byte[]>` serialized
with `FileToBase64Serializer`; ollama-python 0.6.1 `Message.images` holds
`Image` values whose serializer base64-encodes bytes and treats strings as
paths first.
* The Java conversion keeps #1166's tool-call forwarding for assistant
messages alongside the images.
* Verified locally: Ollama module tests 23/23 with spotless; Python Ollama
and chat-message tests pass, with ruff.
</details>
### API
<!-- Does this change touches any public APIs? -->
No new API; it uses the `UnsupportedContentBlockException` /
`UnsupportedContentBlockError` added in #1164.
Compatibility: text-only requests are unchanged. User images used to be
dropped silently; they are now sent. Other media, and media in non-user
messages, now fail instead of being dropped.
### Documentation
<!-- Do not remove this section. Check the proper box only. -->
- [ ] `doc-needed` <!-- Your PR changes impact docs -->
- [ ] `doc-not-needed` <!-- Your PR changes do not impact docs -->
- [x] `doc-included` <!-- Your PR already contains the necessary
documentation updates -->
A multimodal note in the Ollama section of `chat_models.md`, and Ollama
added to the provider-support note.
### Was this patch authored or co-authored using generative AI tooling?
<!-- Do not remove this section. Check the proper box only. -->
- [x] Yes
- [ ] No
If yes, include a `Generated-by: <tool name and version> (<model name and
version>)` line, for example `Generated-by: Claude Code 2.1.226 (Claude Opus
4.6)`, in the commit message so it reaches Git history. Repeat the same line
here for reviewer visibility. See the [ASF generative tooling
guidance](https://www.apache.org/legal/generative-tooling.html).
Generated-by: Claude Code 2.1.259 (Claude Opus 5.5)
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]