Great to hear that! I know the community is busy preparing for the 1.0 release right now, but when you have some bandwidth, or once the RC work is wrapped up, I’d really appreciate it if you or someone from the community could help review this PR.
At the moment, it’s blocked on CI and getting merged. Once this foundational piece lands, I should be able to move on to the next step pretty quickly. Thanks again, and best regards! Shangqing At 2026-09-11 12:35:12, "Giannis Polyzos" <[email protected]> wrote: >It does, and I 100% agree with the approach > >Best, >Giannis > >On Fri, Sep 11, 2026 at 12:47 AM Shangqing Yang <[email protected]> >wrote: > >> Hi Giannis, >> >> >> Thanks a lot for the detailed feedback, these are all very helpful points. >> >> >> I agree that the bounded log scan semantics, retention behavior, bucket >> pruning helper, and the new full KV scan API for primary key tables are all >> worth looking into carefully. I’ll go through these as the design and >> implementation progress, and keep updating the design document accordingly. >> >> >> I also expect the design to continue evolving as Fluss itself evolves and >> exposes more engine-neutral capabilities. Some of them will likely become >> clearer once we start implementing and validating them against the existing >> client APIs. >> >> >> One area where I would especially like to get feedback and, ideally, reach >> some agreement is the overall development and build approach. >> >> >> My current proposal is to develop the Trino connector as a relatively >> isolated module in the Fluss repository, with its own build environment and >> CI, so that it does not change the build requirements or runtime behavior >> of the existing Fluss modules. This would also allow us to develop the >> connector incrementally and merge small, self-contained PRs as we go. >> >> >> I think this is preferable to waiting until the entire connector is >> feature-complete before merging anything. A single PR containing the full >> connector would quickly become very large and difficult to review, mixing >> build infrastructure, Trino SPI integration, metadata handling, scan >> implementations, tests, and optimizations in one change, which would be a >> terrible monster :) >> >> >> With incremental development, each PR can stay focused and independently >> reviewable, and we can adjust the implementation as we get more feedback >> from the community. Of course, merging these foundational or intermediate >> PRs would not mean that the connector is already ready for production use. >> I plan to make the connector status explicit in the documentation until we >> consider it GA. >> >> >> So while I’ll continue refining the implementation details based on >> feedback like yours, I’d particularly appreciate thoughts on whether this >> overall architecture and incremental development model sounds reasonable to >> you and the community. >> >> >> If we agree on this incremental development approach, I’ve already opened >> an initial PR for the basic connector foundation. I’d appreciate any >> reviews and feedback, especially on the module structure and build >> isolation before I continue with the next pieces. >> https://github.com/apache/fluss/pull/4207 >> >> >> Thanks again! this feedback is very useful. >> >> >> Best, >> Shangqing >> >> >> >> >> >> >> >> >> >> >> >> At 2026-09-10 13:31:27, "Giannis Polyzos" <[email protected]> wrote: >> >Hi Shangqing >> >And thank you for driving this. I have already seen requests for a Trino >> >connector, so this is something very much needed that will help fluss open >> >up to more engines. >> >It's a genuinely well-scoped proposal and provides a solid foundation. >> > >> >Some things that come to mind that might be worth looking into: >> > >> >*1. The bounded log split*. >> >The FIP describes a Log split as (bucket, startOffset, stoppingOffset) as >> >if a bounded log read exists, but it doesn't: createBatchScanner on a log >> >table is row-limit bounded, not offset-range bounded. >> >Today you'd build this on the unbounded LogScanner plus >> >Admin.listOffsets(LATEST) to capture stop points and what the Flink >> >connector already does. I'd suggest extracting that into a small >> >engine-neutral "bounded-offset log scanner" in fluss-client (Spark and >> >other engines would benefit too). >> > >> >*2. "Fail rather than silently skip" on log retention needs to be a >> >verified contract.* >> >The default log scanner may auto-reset to the earliest offset when a >> >planned offset has expired, which would quietly violate the guarantee the >> >FIP makes. Worth confirming this behavior and adding some tests to verify >> >this. >> > >> >*3. Bucket pruning needs a small stable helper.* The bucketing function is >> >reusable, but turning predicate values into a bucket id requires the exact >> >canonical key encoding, and those pieces are currently internal. A tiny >> >engine-neutral "key values → bucket id" utility would let the connector >> >prune correctly without duplicating anything. That's the second (and only >> >other) new API I'd call out. >> > >> >*4. PK Table Reads* >> >For the primary key table reads, it might be worth using the newly added >> >api >> > >> https://cwiki.apache.org/confluence/spaces/FLUSS/pages/386272150/FIP-17+Support+Full+KV+Scan+for+Primary+Key+Tables >> > >> > >> >I hope these help a bit.. Let me know your thoughts. >> > >> >Best, >> >Giannis >> > >> >On Wed, Sep 2, 2026 at 7:29 AM Shangqing Yang <[email protected]> >> >wrote: >> > >> >> Hi Fluss community, >> >> >> >> I’d like to start a discussion about adding a native, read-only Trino >> >> connector for Apache Fluss. >> >> >> >> There has already been previous exploration of Trino integration in >> >> FIP-18, #1810, PR #1923, and the related mailing-list discussion. Those >> >> efforts provide valuable prior work, but they also covered several >> >> relatively independent topics at once, including native Fluss reads, >> >> multi-version Trino support, Lakehouse Union Read, Lakehouse readers, >> and >> >> additional pushdowns. >> >> >> >> For this proposal, I’d like to narrow the scope and first establish a >> >> clean native Trino-to-Fluss read path. >> >> >> >> The initial scope includes: >> >> >> >> metadata discovery and type mapping; >> >> bounded reads of non-Lakehouse Log Tables; >> >> current-state reads of non-Lakehouse Primary Key Tables; >> >> projection and predicate-based scan optimizations; >> >> partition and bucket pruning where applicable; >> >> primary-key lookup optimization; >> >> an explicit Trino compatibility and build-isolation policy. >> >> >> >> Lakehouse Union Read, Lakehouse format readers, writes/DDL, and changes >> to >> >> Trino core are explicitly out of scope for this FIP. I think Union Read >> is >> >> an important follow-up topic, and I’d be happy to drive a separate >> proposal >> >> for it after the native connector is established. >> >> >> >> Since I don’t currently have permission to create the FIP page on >> >> Confluence, I’ve prepared the proposal as a Google Doc using FIP-XXX as >> the >> >> temporary identifier: >> >> >> >> Design doc: >> >> >> https://docs.google.com/document/d/1mY-PN8KY1cJSD9YzISuL14A6rEIH1Htj9l99XrLmPk0 >> >> >> >> Issue: https://github.com/apache/fluss/issues/4197 >> >> >> >> I’d especially appreciate feedback on the proposed scope, the Log Table >> >> and Primary Key Table read semantics, and the Fluss Client API boundary. >> >> >> >> Prior work: >> >> >> >> FIP-18: Fluss Trino Engine Support >> >> #1810 >> >> PR #1923 >> >> >> >> Feedback and suggestions are very welcome. >> >> >> >> Thanks, >> >> Shangqing >>
