It does, and I 100% agree with the approach Best, Giannis
On Fri, Sep 11, 2026 at 12:47 AM Shangqing Yang <[email protected]> wrote: > Hi Giannis, > > > Thanks a lot for the detailed feedback, these are all very helpful points. > > > I agree that the bounded log scan semantics, retention behavior, bucket > pruning helper, and the new full KV scan API for primary key tables are all > worth looking into carefully. I’ll go through these as the design and > implementation progress, and keep updating the design document accordingly. > > > I also expect the design to continue evolving as Fluss itself evolves and > exposes more engine-neutral capabilities. Some of them will likely become > clearer once we start implementing and validating them against the existing > client APIs. > > > One area where I would especially like to get feedback and, ideally, reach > some agreement is the overall development and build approach. > > > My current proposal is to develop the Trino connector as a relatively > isolated module in the Fluss repository, with its own build environment and > CI, so that it does not change the build requirements or runtime behavior > of the existing Fluss modules. This would also allow us to develop the > connector incrementally and merge small, self-contained PRs as we go. > > > I think this is preferable to waiting until the entire connector is > feature-complete before merging anything. A single PR containing the full > connector would quickly become very large and difficult to review, mixing > build infrastructure, Trino SPI integration, metadata handling, scan > implementations, tests, and optimizations in one change, which would be a > terrible monster :) > > > With incremental development, each PR can stay focused and independently > reviewable, and we can adjust the implementation as we get more feedback > from the community. Of course, merging these foundational or intermediate > PRs would not mean that the connector is already ready for production use. > I plan to make the connector status explicit in the documentation until we > consider it GA. > > > So while I’ll continue refining the implementation details based on > feedback like yours, I’d particularly appreciate thoughts on whether this > overall architecture and incremental development model sounds reasonable to > you and the community. > > > If we agree on this incremental development approach, I’ve already opened > an initial PR for the basic connector foundation. I’d appreciate any > reviews and feedback, especially on the module structure and build > isolation before I continue with the next pieces. > https://github.com/apache/fluss/pull/4207 > > > Thanks again! this feedback is very useful. > > > Best, > Shangqing > > > > > > > > > > > > At 2026-09-10 13:31:27, "Giannis Polyzos" <[email protected]> wrote: > >Hi Shangqing > >And thank you for driving this. I have already seen requests for a Trino > >connector, so this is something very much needed that will help fluss open > >up to more engines. > >It's a genuinely well-scoped proposal and provides a solid foundation. > > > >Some things that come to mind that might be worth looking into: > > > >*1. The bounded log split*. > >The FIP describes a Log split as (bucket, startOffset, stoppingOffset) as > >if a bounded log read exists, but it doesn't: createBatchScanner on a log > >table is row-limit bounded, not offset-range bounded. > >Today you'd build this on the unbounded LogScanner plus > >Admin.listOffsets(LATEST) to capture stop points and what the Flink > >connector already does. I'd suggest extracting that into a small > >engine-neutral "bounded-offset log scanner" in fluss-client (Spark and > >other engines would benefit too). > > > >*2. "Fail rather than silently skip" on log retention needs to be a > >verified contract.* > >The default log scanner may auto-reset to the earliest offset when a > >planned offset has expired, which would quietly violate the guarantee the > >FIP makes. Worth confirming this behavior and adding some tests to verify > >this. > > > >*3. Bucket pruning needs a small stable helper.* The bucketing function is > >reusable, but turning predicate values into a bucket id requires the exact > >canonical key encoding, and those pieces are currently internal. A tiny > >engine-neutral "key values → bucket id" utility would let the connector > >prune correctly without duplicating anything. That's the second (and only > >other) new API I'd call out. > > > >*4. PK Table Reads* > >For the primary key table reads, it might be worth using the newly added > >api > > > https://cwiki.apache.org/confluence/spaces/FLUSS/pages/386272150/FIP-17+Support+Full+KV+Scan+for+Primary+Key+Tables > > > > > >I hope these help a bit.. Let me know your thoughts. > > > >Best, > >Giannis > > > >On Wed, Sep 2, 2026 at 7:29 AM Shangqing Yang <[email protected]> > >wrote: > > > >> Hi Fluss community, > >> > >> I’d like to start a discussion about adding a native, read-only Trino > >> connector for Apache Fluss. > >> > >> There has already been previous exploration of Trino integration in > >> FIP-18, #1810, PR #1923, and the related mailing-list discussion. Those > >> efforts provide valuable prior work, but they also covered several > >> relatively independent topics at once, including native Fluss reads, > >> multi-version Trino support, Lakehouse Union Read, Lakehouse readers, > and > >> additional pushdowns. > >> > >> For this proposal, I’d like to narrow the scope and first establish a > >> clean native Trino-to-Fluss read path. > >> > >> The initial scope includes: > >> > >> metadata discovery and type mapping; > >> bounded reads of non-Lakehouse Log Tables; > >> current-state reads of non-Lakehouse Primary Key Tables; > >> projection and predicate-based scan optimizations; > >> partition and bucket pruning where applicable; > >> primary-key lookup optimization; > >> an explicit Trino compatibility and build-isolation policy. > >> > >> Lakehouse Union Read, Lakehouse format readers, writes/DDL, and changes > to > >> Trino core are explicitly out of scope for this FIP. I think Union Read > is > >> an important follow-up topic, and I’d be happy to drive a separate > proposal > >> for it after the native connector is established. > >> > >> Since I don’t currently have permission to create the FIP page on > >> Confluence, I’ve prepared the proposal as a Google Doc using FIP-XXX as > the > >> temporary identifier: > >> > >> Design doc: > >> > https://docs.google.com/document/d/1mY-PN8KY1cJSD9YzISuL14A6rEIH1Htj9l99XrLmPk0 > >> > >> Issue: https://github.com/apache/fluss/issues/4197 > >> > >> I’d especially appreciate feedback on the proposed scope, the Log Table > >> and Primary Key Table read semantics, and the Fluss Client API boundary. > >> > >> Prior work: > >> > >> FIP-18: Fluss Trino Engine Support > >> #1810 > >> PR #1923 > >> > >> Feedback and suggestions are very welcome. > >> > >> Thanks, > >> Shangqing >
