Hi Hongshun,

Good point — fluss-rust is currently at 0.1.0 with master tracking 0.2.0, so 
"0.3.0" was just our shorthand here. Once FIP-40 lands and fluss-rust moves 
into the fluss repo, it would naturally follow fluss versioning — likely the 
release after 1.0. I'll update the framing to reflect that.


Best regards,
Junbo Wang

> 2026年7月3日 14:16,Hongshun Wang <[email protected]> 写道:
> 
> Hi Junbo,
> What's 0.3.0? Since fluss-rust will be merged into fluss repo, maybe fluss
> release-1.0?
> 
> Best,
> Hongshun
> 
> On Fri, Jul 3, 2026 at 11:27 AM Junbo Wang <[email protected]> wrote:
> 
>> Thanks Forward and Hongshun for the thoughtful input.
>> 
>> On ForwardXu's question about the target release — I've been thinking a
>> bit about the 0.3.0 scope, and would like to share my personal take, mostly
>> as a starting point for discussion.
>> 
>> A possible 0.3.0 scope
>> 
>> FIP-40: consolidate fluss-rust into apache/fluss. thanks to Anton for
>> already opening the PR-3401 <https://github.com/apache/fluss/pull/3401>
>> DataFusion integration adapter (Theme 1), as a fluss-rust submodule.
>> Lake + log union read, Paimon-first (Theme 2), as a separate submodule so
>> lake-side dependencies stay isolated from the core. S
>> erver-side filter pushdown (Theme 3) — would be great to have if capacity
>> allows, but I'd suggest keeping it optional rather than a hard requirement
>> for 0.3.0.
>> 
>> Gateway — perhaps outside 0.3.0
>> 
>> Once fluss-rust is consolidated into apache/fluss, it might make sense for
>> the gateway to live as its own module depending on fluss-rust, rather than
>> being tied to the 0.3.0 release train. That would also give the auth/authz
>> questions Hongshun raised some room to be designed in a dedicated FIP.
>> 
>> Just my personal thinking — happy to adjust based on what others feel is
>> right.
>> 
>> 
>> Best regards,
>> Junbo Wang
>> 
>>> 2026年7月2日 11:47,Forward Xu <[email protected]> 写道:
>>> 
>>> Hi Anton,
>>> 
>>> Thanks for putting this together, and thanks to everyone who landed the
>>> previous roadmap. Framing the next phase around "fluss-rust as a
>>> first-class analytical query surface" makes a lot of sense to me — it
>>> builds naturally on the analytical primitives we just shipped. A few
>>> thoughts, roughly following your structure.
>>> 
>>> *Theme 1 — DataFusion integration adapter [MUST-HAVE]* +1. I strongly
>> agree
>>> it should be an *adapter over the Rust core*, not a bundled engine.
>> Keeping
>>> fluss-datafusion as a standalone crate exposing TableProvider +
>>> CatalogProvider/SchemaProvider keeps the dependency direction clean and
>>> lets any DataFusion-based consumer (or the gateway) opt in. Mapping
>>> projection/filter/limit pushdown onto the existing access paths (PK
>>> equality → lookup, bucket-key prefix → prefix lookup, LIMIT → bounded
>> scan,
>>> else log scan) is the right approach. One thing worth nailing down early
>> is
>>> the exactness contract we report back to DataFusion (
>>> TableProviderFilterPushDown::Exact vs Inexact), since getting that wrong
>>> silently drops or double-applies filters.
>>> 
>>> *Theme 2 — Lake + log union read (Paimon-first) [MUST-HAVE]* +1, and
>>> Paimon-first is the right call — reusing paimon-rust by wrapping an
>>> existing reader is low-risk, and the lake-snapshot-records-log-offset
>>> design gives us a clean stitch point. Starting bounded-batch and
>> deferring
>>> streaming also de-risks correctness of the log-offset boundary before we
>>> add continuous reads. The main correctness edge I'd like to see covered
>> is
>>> the boundary semantics for PK tables when applying newer log records over
>>> the lake snapshot (dedup/ordering at the tiered offset).
>>> 
>>> *Theme 3 — Foundations / server-side filter pushdown [STRETCH]* Agree
>> this
>>> is a stretch, but it's the item with the highest leverage since it prunes
>>> both lake and log scans and unblocks the gateway's filter pushdown.
>>> Cross-checking against the Java PredicateConverter semantics is important
>>> so both clients behave identically. If capacity allows, I'd lean toward
>>> pulling at least the simple value-predicate cases forward, since Themes 1
>>> and 4 both benefit.
>>> 
>>> *Theme 4 — Gateway [STRETCH]* The "REST is the no-regret floor, SQL is
>> the
>>> ceiling deferred until Themes 1–2" framing is a good way to think about
>> it.
>>> Keeping it a thin, stateless HTTP frontend over the existing client
>>> primitives (one endpoint per primitive) sounds right and keeps scope
>>> contained.
>>> 
>>> *Open questions*
>>> 
>>>  - *Themes 1–2 must-have, rest stretch — reasonable?* Yes, ordering
>> looks
>>>  right to me. Nothing obviously missing.
>>>  - *Paimon first, Iceberg later?* Agree, provided the union-read path is
>>>  written against a lake abstraction so Iceberg/Lance slot in the same
>> way
>>>  rather than requiring a rewrite.
>>>  - *Pin one arrow/DataFusion version across core, union-read/lake, and
>>>  DF?* Strong +1 — we should pin a single Arrow/DataFusion version
>>>  workspace-wide; version skew between these crates is a common and
>> painful
>>>  source of breakage.
>>>  - *Separate crates: lake kernel, DataFusion adapter, gateway?* +1 to
>>>  separate crates. It matches the "adapter, not bundled engine"
>> principle and
>>>  keeps the DataFusion dependency out of the core for consumers that
>> don't
>>>  need it.
>>> 
>>> One meta-question: do we have a rough sense of the target release
>> (0.3.0?)
>>> these must-haves land in, so we can size the Theme 3/4 stretch work
>>> accordingly?
>>> 
>>> Thanks again — happy to help on the DataFusion adapter side.
>>> 
>>> Best,
>>> 
>>> ForwardXu
>>> 
>>> Junbo Wang <[email protected]> 于2026年7月1日周三 22:41写道:
>>> 
>>>>> I'd also like us to focus on closing the gap with the Java client on
>>>> predicate pushdown.
>>>> Agreed — predicate pushdown is a great optimization, and the read path
>>>> will benefit significantly from it.
>>>> 
>>>> 
>>>> 
>>>> One thing that was missing in upstream paimon-rust was an
>>>> object-store-safe existence check in the filesystem catalog. It checked
>>>> whether database/table directories “exist”, which works on local
>>>> filesystems but can fail on S3/OSS because those directories are just
>>>> prefixes, not real objects. I forked it to fix that, so lake reads can
>>>> reliably open Paimon tables from object storage.
>>>> 
>>>> Thanks again, Anton, for putting this roadmap together! I think we could
>>>> share it at the July 3rd community meeting. I've also been working on
>> some
>>>> designs around lake kernel read and the Fluss Gateway REST API — happy
>> to
>>>> discuss those as well if time allows.
>>>> 
>>>> 
>>>> Best regards,
>>>> Junbo Wang
>>>> 
>>>>> 2026年6月25日 07:23,Anton Borisov <[email protected]> 写道:
>>>>> 
>>>>> Hi Yuxia and Junbo,
>>>>> 
>>>>> Yuxia, on PK changelog reads across the bindings: I have already
>>>>> prepared record-mode CDC for Rust, Python and C++. It's small, and
>>>>> the PK union read in Theme 2 needs it anyway.
>>>>> 
>>>>> Junbo, I looked through your prototype branch. It is clear that a lot
>>>>> of Themes 1, 2 and 4 already exist there as working code and it does
>>>>> seem to answer two of the open questions: the three-way crate split
>>>>> (lake kernel / adapter / gateway), and pinning a single
>> Arrow/DataFusion
>>>>> version at the workspace level. +1 to both.
>>>>> 
>>>>> Two separate FIPs with detailed design for REST/SQL gateway parts -
>>>> makes sense.
>>>>> 
>>>>> One thing I noticed and want to discuss:
>>>>> The "lake read" currently depends on a fork of paimon-rust, can you
>> share
>>>>> what was missing in paimon-rust?
>>>>> 
>>>>> I'd also like us to focus on closing the gap with the Java client on
>>>>> predicate pushdown.
>>>>> In the interim we can push predicates into the lake (Paimon already
>>>>> supports it) and apply them as a filter pass over the real-time log
>>>>> tail.
>>>>> 
>>>>> Good to hear you want to drive this forward, given the prototype, that
>>>>> makes sense to me. I am happy to review the design and following PRs.
>>>>> 
>>>>> Looking forward to the write-up with your findings.
>>>>> 
>>>>> -- Anton
>>>>> 
>>>>> вт, 23 июн. 2026 г. в 15:10, Junbo Wang <[email protected]>:
>>>>>> 
>>>>>> Thanks Anton for kicking this off — funnily enough, I'd been arriving
>>>> at very similar conclusions while prototyping on my side, so let me
>> share a
>>>> bit of what came out of those experiments.
>>>>>> 
>>>>>>> THEME 1: DATAFUSION INTEGRATION ADAPTER [MUST-HAVE]
>>>>>> 
>>>>>> Agree. Having the DataFusion integration live in fluss-rust would make
>>>> it reusable across more downstream projects. One thing we might want to
>>>> keep in mind: it would probably be worth aligning the DataFusion and
>> Arrow
>>>> versions with pg-datafusion so the whole stack stays on a consistent
>> set of
>>>> versions.
>>>>>> 
>>>>>>> THEME 2: LAKE + LOG UNION READ (PAIMON-FIRST) [MUST-HAVE]
>>>>>> 
>>>>>> Agree. I'd lean toward putting this in a separate fluss-lake module
>>>> that exposes a client capable of doing the union read over the lake and
>> the
>>>> Fluss log. That way the lake-side dependencies stay isolated in their
>> own
>>>> module and don't bleed into the core fluss-rust code. One thing that
>> might
>>>> be worth flagging here: the lake side already supports predicate
>> pushdown,
>>>> while Fluss doesn't yet — so we'd probably think about how the union
>> read
>>>> handles that asymmetry in the interim.
>>>>>> 
>>>>>> THEME 4: GATEWAY [STRETCH] — A thin HTTP frontend over the client
>>>>>> 
>>>>>> I wonder if we could expand on this one a bit. My suggestion would be
>>>> to split the SQL REST surface and the plain REST API into two separate
>>>> FIPs. A read/write REST API on its own — somewhat in the spirit of
>>>> kafka-rest, covering reads, writes, and metadata management — would
>> already
>>>> be a quick win for broadening Fluss's data ingestion/access reach, and
>> it
>>>> doesn't depend on anything else. The SQL REST surface, on the other
>> hand,
>>>> has to wait on Themes 1–2, so it might make sense to split them and ship
>>>> the plain REST API first.
>>>>>> 
>>>>>> 
>>>>>> 
>>>>>> Also, it might make sense to have the gateway as a separate crate
>>>> inside the Fluss project itself — especially since fluss-rust will be
>>>> moving into the Fluss repo anyway.
>>>>>> 
>>>>>> On a personal note, this is something I'm really excited about and
>>>> would love to help push forward. I'll share a write-up of what came out
>> of
>>>> my demo experiments with the community over the next few weeks.
>>>>>> 
>>>>>> 
>>>>>> Best regards,
>>>>>> Junbo Wang
>>>>>> 
>>>>>>> 2026年6月23日 20:55,Yuxia Luo <[email protected]> 写道:
>>>>>>> 
>>>>>>> Thanks for putting this together. +1 on the overall framing. My
>>>> answers to the open questions, plus one addition:
>>>>>>> 
>>>>>>> Open questions:
>>>>>>> - Themes 1-2 as must-have, rest stretch: agree, the ordering makes
>>>> sense to me.
>>>>>>> - Paimon first, Iceberg later: +1. Reusing paimon-rust is the
>>>> pragmatic path, and the union-read design generalizes to Iceberg the
>> same
>>>> way later.
>>>>>>> - Pin one arrow/DataFusion version across core, union-read/lake, and
>>>> DF: yes, we should pin a single version. Letting these drift will cause
>>>> painful Arrow type/ABI mismatches across the crate boundary, so a shared
>>>> pinned version (bumped deliberately, not per-crate) is worth the
>> discipline.
>>>>>>> - Separate crates (lake kernel, DataFusion adapter, gateway): +1 on
>>>> separate crates. It keeps the dependency surface clean - the adapter and
>>>> gateway are optional consumers, and the lake kernel shouldn't drag
>>>> DataFusion into anyone who only needs the core read path.
>>>>>>> 
>>>>>>> One addition I'd like to propose for the scope:
>>>>>>> PK-table CHANGELOG read, and expose it across the bindings -
>>>> specifically C++ and Python - so non-Rust CDC consumers can subscribe.
>>>>>>> 
>>>>>>> On 2026/06/22 14:49:52 Anton Borisov wrote:
>>>>>>>> Hi all,
>>>>>>>> 
>>>>>>>> With the previous roadmap wrapping up - complex types
>>>>>>>> (Array/Row/Map/nesting), limit and prefix scan, schema-aware KV
>>>>>>>> decoding, the metrics framework, and write optimisations, thanks to
>>>>>>>> everyone who contributed and reviewed, I'd like to open discussion
>> on
>>>>>>>> the roadmap.
>>>>>>>> 
>>>>>>>> I propose framing around a single goal: make fluss-rust a
>> first-class
>>>>>>>> analytical query surface - a Rust-native SQL/DataFrame path over
>> Fluss
>>>>>>>> (DataFusion, and through it Polars/DuckDB/gateways), reading the
>> lake
>>>>>>>> tier at scale, built on the analytical primitives we just landed.
>>>>>>>> 
>>>>>>>> As before, I grouped the items into themes with an initial
>> must-have /
>>>>>>>> stretch positioning. Please push back where you disagree.
>>>>>>>> 
>>>>>>>> THEME 1: DATAFUSION INTEGRATION ADAPTER [MUST-HAVE]
>>>>>>>> A standalone fluss-datafusion crate exposing TableProvider + Catalog
>>>>>>>> over the existing client, so SELECT ... FROM . works from any
>>>>>>>> DataFusion-based engine. It should be framed as an integration
>> adapter
>>>>>>>> over the Rust core, not a bundled engine. A gateway (FIP-32) or any
>>>>>>>> analytical consumer can use it or not.
>>>>>>>> 
>>>>>>>> Scope: TableProvider + CatalogProvider/SchemaProvider backed by the
>>>>>>>> metadata path.
>>>>>>>> 
>>>>>>>> Map DataFusion's projection/filter/limit pushdown onto the client's
>>>>>>>> existing access paths: a full primary-key equality becomes a
>> lookup, a
>>>>>>>> bucket-key prefix becomes a prefix lookup, LIMIT becomes a bounded
>>>>>>>> scan, otherwise a log scan. Filters an access path fully satisfies
>> are
>>>>>>>> reported exact, the rest are left for DataFusion to apply
>> (server-side
>>>>>>>> filter pushdown is Theme 3).
>>>>>>>> 
>>>>>>>> THEME 2: LAKE + LOG UNION READ (PAIMON-FIRST) [MUST-HAVE]
>>>>>>>> A client-side read that presents a lake-enabled table as one table.
>>>>>>>> Its history is tiered to the lake (Paimon) and recent writes are
>> still
>>>>>>>> in the Fluss log, the lake snapshot records the log offset it
>> covers,
>>>>>>>> so we read the lake up to that offset and the Fluss log past it,
>> then
>>>>>>>> combine.
>>>>>>>> Scope:
>>>>>>>> - Per bucket: the lake snapshot plus the Fluss log past the tiered
>>>> offset.
>>>>>>>> - Log tables: concatenate the two. PK tables: apply the newer log
>>>>>>>> records to get current state.
>>>>>>>> - Paimon first, reusing paimon-rust, Iceberg/Lance later the same
>>>>>>>> way. The DataFusion adapter reads lake-enabled tables through this
>>>>>>>> path.
>>>>>>>> - Bounded batch first, streaming later
>>>>>>>> 
>>>>>>>> Reusing paimon-rust means wrapping an existing reader, so it's
>> rather
>>>>>>>> feasible and straightforward.
>>>>>>>> 
>>>>>>>> THEME 3: FOUNDATIONS [STRETCH]
>>>>>>>> Server-side value-predicate (filter) pushdown. The proto field
>> exists,
>>>>>>>> but the client currently sends none. This prunes lake/log scans and
>>>>>>>> unblocks the gateway's filter pushdown (the open question I raised
>> for
>>>>>>>> 0.2.0), cross-checked against the Java PredicateConverter semantics.
>>>>>>>> 
>>>>>>>> THEME 4: GATEWAY [STRETCH]
>>>>>>>> A thin HTTP frontend over the client we already have  independent of
>>>>>>>> Themes 1–3. It's how non-Rust/non-SQL clients (TS/JS, microservices,
>>>>>>>> curl) write to Fluss and do simple key access.
>>>>>>>> - Write: upsert/append, delete, create/drop tables, metadata. The
>>>>>>>> mature client write path essentially.
>>>>>>>> - Read: point lookup, prefix lookup, bounded log / CDC poll - one
>>>>>>>> endpoint per existing primitive, bounded and stateless.
>>>>>>>> The tradeoff: no joins, filters, or aggregations - that's the SQL
>>>>>>>> surface, and it needs the adapter and probably will be an ongoing
>>>>>>>> effort to optimise later.
>>>>>>>> - If Themes 1–2 ship, we unlock SQL frontends on top - PostgreSQL
>>>>>>>> via datafusion-postgres, as one example
>>>>>>>> In short: REST is the no-regret floor, SQL is the ceiling deferred
>>>>>>>> until Themes 1–2.
>>>>>>>> 
>>>>>>>> PARALLEL TRACK: ELIXIR + BINDING EXPOSURE
>>>>>>>> Continue Elixir parity and expose newly-landed primitives (prefix
>>>>>>>> lookup, etc.) across Python / C++ / Elixir.
>>>>>>>> 
>>>>>>>> GOVERNANCE TRACK
>>>>>>>> FIP-40: Consolidate apache/fluss-rust into apache/fluss - PR is
>> ready,
>>>>>>>> looks pretty good, waiting for a good moment to rebase/merge.
>>>>>>>> 
>>>>>>>> OPEN QUESTIONS
>>>>>>>> - Themes 1–2 must-have, the rest stretch. Reasonable? Anything
>>>>>>>> missing or mis-ordered?
>>>>>>>> - Paimon first, Iceberg later?
>>>>>>>> - Pin one arrow/DataFusion version across core, union-read/lake, and
>>>> DF?
>>>>>>>> - Separate crates: lake kernel, DataFusion adapter, gateway?
>>>>>>>> 
>>>>>>>> Looking forward to feedback.
>>>>>>>> 
>>>>>>>> -- Anton
>>>>>>>> 
>>>>>> 
>>>> 
>>>> 
>> 
>> 

Reply via email to