Hi Sagar,

And thanks for tackling this.. Currently everyone is super busy with the
release coming these days. In case no one has time to take a look within
then next few days, allow me a week or two when I will be fully back to
look into this, because it’s something I’m excited about.

Apologies for the delays and do know your work is highly appreciated πŸ™

We can also sync on the upcoming Fluss community call.

Best,
Giannis

On Sat, 22 Aug 2026 at 4:16β€―AM, Sagar <[email protected]> wrote:

> Hi ,
>
> If someone got a chance to look at the poc, Please let me know if this is
> at a stage when I can create an fip for this.
>
> Sagar.
>
> On Sat, 15 Aug 2026 at 9:54β€―PM, Sagar <[email protected]> wrote:
>
> > Thanks Mehul!
> >
> > I Wanted to share an update on the VECTOR data type for Fluss. This PR
> has
> > an e2e POC for the same:
> >
> > https://github.com/apache/fluss/pull/4004.
> >
> > It mainly contains an IT for Lance Vector Tiering
> > (LanceVectorTieringITCase) with 2 tests: testVectorTiering and
> > testNullVectorTiering. I am seeing some build errors on fluss-spark but
> the
> > IT seems to work locally.  At a high level, it:
> >
> >    - Writes records with vector embeddings (VECTOR(dim)) into a Fluss
> >    table.
> >    - A Flink background job reads those vector records from Fluss and
> >    writes them into Lance storage.
> >    - The test opens the resulting Lance dataset directly and checks that
> >    every vector, float value, row count, and null field made it through
> >    accurately without corruption or data loss.
> >
> > While this seems to work, there are a couple of things worth calling out:
> >
> > 1) Flink SQL Limitation: Flink SQL doesn't have a native VECTOR data type
> > (it treats embeddings as ARRAY<FLOAT>), so vector columns aren't exposed
> as
> > a distinct VECTOR type in SQL queries. Under the hood, Fluss models this
> as
> > a dedicated VECTOR(dim) type mapped directly to Arrow's
> > FixedSizeListVector<Float32>. This avoids the extra overhead of
> > variable-sized lists and aligns 1:1 with Lance's native vector storage
> > format.
> >
> > 2) Multi-Dimensional Array Limitation
> > Currently, we only support 1D fixed-size float vectors. Multi-dimensional
> > arrays (like matrices or nested arrays ARRAY<ARRAY<FLOAT>>) are not
> > supported for vector tiering yet.
> >
> > To support that, we would need more work across schema conversion to
> Lance
> > multi-level list format and even more things on Arrow converters. We can
> > take a look at it later i think.
> >
> > Please take a look and see if it is at a point where we can introduce a
> > FIP. I think it would be similar to this doc:
> >
> >
> >
> https://docs.google.com/document/d/1idmsgMLjYScgYj-l_ABD7rwvNbDvoaq8bbMgWcgZW20/edit?tab=t.0
> >
> > Sagar.
> >
> >
> > On Sat, Jul 4, 2026 at 5:45β€―PM Mehul Batra <[email protected]>
> > wrote:
> >
> >> Hi Sagar,
> >>
> >> Thanks for the detailed findings and document, really appreciate the
> work
> >> here.
> >>
> >> I think this is the right direction  as discussed on the community call,
> >> let’s keep phase one focused on just the vector data type. Excited to
> see
> >> the POC take shape, and a FIP to discuss the findings sounds like a
> great
> >> next step.
> >>
> >> A few things to keep in mind from our thread & discussions:
> >>
> >> β€’ Fixed dimension + fixed element type at schema time (VECTOR(1536))
> >>
> >> β€’ Arrow-native layout (FixedSizeList<Float32>, zero-copy with
> >> Lance/Paimon)
> >>
> >> β€’ Support FLOAT32, FLOAT16, INT8 for quantization from day one so we
> avoid
> >> a breaking change later.
> >>
> >> Best Regards,
> >>
> >> Mehul Batra
> >> On Sun, Jun 28, 2026 at 10:21β€―AM Sagar <[email protected]>
> wrote:
> >>
> >> > Hi,
> >> >
> >> > Following up again to see if there are any further comments or
> feedback.
> >> >
> >> > Sagar.
> >> >
> >> > On Tue, 16 Jun 2026 at 11:09β€―PM, Sagar <[email protected]>
> >> wrote:
> >> >
> >> > > Hi Jark
> >> > >
> >> > > Thanks for the detailed feedback! Please find my responses:
> >> > >
> >> > >
> >> > > *Point 1*
> >> > >
> >> > > *Property-to-API mapping and unexposed parameters* β€” Added a mapping
> >> > > table to Public Interfaces covering all three API calls. For
> >> parameters
> >> > > Fluss doesn't expose: replace is fixed True on both create_index()
> and
> >> > > create_fts_index() β€” any other value would break the idempotency the
> >> > > state machine relies on. use_tantivy is fixed False (see below).
> >> > num_bits,
> >> > > delete_unverified, and retrain are not exposed; LanceDB defaults
> >> apply.
> >> > > Since we route through JNI to lance-core (discussed below) rather
> than
> >> > the
> >> > > Python client, parameter semantics are identical to the Python
> >> > equivalents
> >> > > shown in the example.
> >> > >
> >> > > *FTS path* β€” Fluss targets the native FTS path (use_tantivy=False),
> >> now
> >> > > the LanceDB upstream default. The legacy Tantivy path is not
> >> supported;
> >> > it
> >> > > differs in both parameter surface and on-disk format. This is fixed
> at
> >> > the
> >> > > JNI layer.
> >> > >
> >> > > *Default divergences* β€” Checked against the LanceDB docs [1]: all
> FTS
> >> > > defaults in the FIP are identical to LanceDB's defaults, so no
> >> rationale
> >> > is
> >> > > needed there. The only divergences are on the vector side:
> >> > ef_construction
> >> > > (Fluss: 150, LanceDB: 300) and lance.index.m=(none), which maps to
> >> > > LanceDB's hardcoded 20. Both are now documented with rationale.
> >> > >
> >> > > *Side-by-side example* β€” Added below the existing full-configuration
> >> SQL
> >> > > block.
> >> > >
> >> > > *Point 2 β€” Execution model, LanceDB embedding, and horizontal
> scaling*
> >> > >
> >> > > *Where does the committer run?*
> >> > >
> >> > > Inside the Tiering Service worker, consistent with FIP-5. Modified
> the
> >> > FIP
> >> > > to update this
> >> > >
> >> > > *How is LanceDB embedded in the JVM?*
> >> > >
> >> > > com.lancedb:lance-core is a first-party JNI binding from the
> >> > > lance-format/lance monorepo β€” already a dependency in Fluss's
> existing
> >> > > Lance integration from FIP-5, not a new one introduced here.
> >> > >
> >> > > I verified the published 0.39.0 JAR directly. createIndex and
> >> listIndexes
> >> > > are present. The one gap is optimizeIndices β€” needed to fold newly
> >> > > written rows into existing indices after each tiering cycle. The JNI
> >> > > pattern is established in the codebase by nativeCreateIndex;
> >> contributing
> >> > > nativeOptimizeIndices is a single function addition in
> >> > > java/lance-jni/src/dataset.rs with a corresponding method pair in
> >> > > Dataset.java. This is a committed prerequisite of the FIP-44
> >> > > implementation. No sidecar, no subprocess. FIP is updated with this
> >> > detail.
> >> > >
> >> > > Similarly, the APIs to add an FTS index also seem missing in the jni
> >> > > binding. We will need to add those as well.
> >> > >
> >> > > *How is index work distributed?*
> >> > >
> >> > > Per-table, scoped to the committer owning that table. Horizontal
> >> scaling
> >> > > is at table granularity.
> >> > >
> >> > > The single-table pinning concern is real but bounded: createIndex is
> >> > > non-blocking β€” the build runs async inside LanceDB's Tokio runtime,
> >> the
> >> > > committer thread is released immediately and polls listIndexes() on
> >> > > subsequent timer fires. The build itself is internally
> multi-threaded.
> >> > The
> >> > > constraint is cross-JVM-process parallelism, not single-threading. A
> >> > *Scaling
> >> > > Constraints* note will be added to the FIP. Coordinator-assigned
> index
> >> > > builds are a reasonable future extension but out of scope here. This
> >> is
> >> > > also added to the FIP.
> >> > >
> >> > >
> >> > > *3. Configuration drift after the table exists*
> >> > >
> >> > > For the initial scope of FIP-44, we will take the *'reject at DDL
> >> time'*
> >> > > approach.
> >> > >
> >> > > Index configurations will be treated as immutable once the index
> state
> >> > > enters IN_PROGRESS or COMPLETED. If a user attempts to modify
> >> properties
> >> > > like lance.index.type, metric, or num_partitions via ALTER TABLE,
> the
> >> DDL
> >> > > validator will reject it.
> >> > >
> >> > > *Rationale:* This keeps the FIP-44 state machine strictly linear
> >> (ABSENT
> >> > > β†’ IN_PROGRESS β†’ COMPLETED). It avoids the complexities of modeling
> >> > > PENDING_REBUILD states and protects Tiering workers from
> accidentally
> >> > > triggering massive background rebuilds due to a simple property
> tweak.
> >> > > Declarative background rebuilds for config drift can be tackled in a
> >> > future
> >> > > FIP. I will update the document to explicitly state this constraint
> >> > >
> >> > > Let me know what you think!
> >> > >
> >> > >
> >> > > Sagar.
> >> > >
> >> > > [1]:
> https://docs.lancedb.com/search/full-text-search#advanced-usage
> >> > >
> >> > > On Sun, May 31, 2026 at 10:34β€―AM Jark Wu <[email protected]> wrote:
> >> > >
> >> > >> Hi Sagar,
> >> > >>
> >> > >> Thanks for the detailed FIP. Three comments below.
> >> > >>
> >> > >> ## 1. Public-interface docs need a mapping and a worked example
> >> > >>
> >> > >> The `lance.*` properties currently stand alone in the FIP β€” to
> >> > >> understand any of them, a reader has to cross-reference the LanceDB
> >> > >> docs. I'd like the FIP to add three things to the public-interface
> >> > >> section:
> >> > >>
> >> > >> - An explicit table mapping each Fluss property to the LanceDB API
> >> > >> call and parameter it maps to (e.g. `lance.index.type` β†’
> >> > >> `Table.create_index(index_type=...)`).
> >> > >> - An explicit mapping of each Fluss default to the corresponding
> >> > >> LanceDB default, with rationale for any deliberate divergence.
> >> > >> Skimming the FIP, several `lance.fts.*` defaults look like they
> >> differ
> >> > >> from LanceDB upstream defaults (e.g. `stem`, `remove_stop_words`,
> >> > >> `ascii_folding`), and `lance.index.m`'s `(none)` effectively means
> >> > >> LanceDB's hardcoded `20`. The reasons aren't stated.
> >> > >> - A side-by-side example showing the same index expressed as (a) a
> >> > >> Fluss `CREATE TABLE ... WITH (...)` statement, and (b) the
> equivalent
> >> > >> LanceDB Python call. That makes the abstraction concrete for both
> >> > >> reviewers and future users.
> >> > >>
> >> > >> Two specific things worth pinning down while you're in there:
> >> > >>
> >> > >> - `create_fts_index` has a legacy Tantivy path and a newer native
> FTS
> >> > >> path (`use_tantivy=False`, now the upstream default). Which one is
> >> the
> >> > >> FIP targeting? The parameter surface and on-disk format both
> differ.
> >> > >> - `Table.create_index` and `Table.optimize` have additional
> >> parameters
> >> > >> (`replace`, `num_bits`, `delete_unverified`, `retrain`, …) that
> >> aren't
> >> > >> currently mapped. Either include them or explain why they're
> >> > >> deliberately hidden β€” `replace` in particular matters because the
> >> > >> state machine relies on `create_index` being idempotent, which is
> >> only
> >> > >> true with `replace=True`.
> >> > >>
> >> > >> ## 2. Who builds the index? Execution model and horizontal scaling
> >> > >>
> >> > >> The FIP assigns the index lifecycle to the `LanceLakeCommitter`,
> but
> >> > >> the deeper execution-model question is not yet answered:
> >> > >>
> >> > >> - Where does the committer (and therefore `create_index()` /
> >> > >> `optimize()`) physically run? My reading of FIP-5 is that the
> >> > >> committer lives inside the Tiering Service workers. Is that the
> >> intent
> >> > >> here?
> >> > >>
> >> > >> - If so, the Tiering Service now has to **embed LanceDB**. LanceDB
> is
> >> > >> a Rust core with Python and Node bindings β€” there is no first-party
> >> > >> Java client today. How is it embedded into the JVM-based tiering
> >> > >> worker? JNI over the Rust core? A sidecar subprocess? Something
> else?
> >> > >> This is a non-trivial dependency to take on and deserves explicit
> >> > >> discussion in the FIP.
> >> > >>
> >> > >> - How is index work **distributed** across Tiering Service workers?
> >> > >> Per-table affinity? Coordinator-assigned? With a single large table
> >> > >> whose one heavy index takes hours to build, does the work pin to
> one
> >> > >> worker, or can it be split? If the asynchronous build effectively
> >> runs
> >> > >> in-process inside the worker that initiated it, then horizontal
> >> > >> scaling is per-table at best.
> >> > >>
> >> > >>
> >> > >> ## 3. Configuration drift after the table exists
> >> > >>
> >> > >> What happens if a user changes `lance.index.type` (or `metric`,
> >> > >> `num_partitions`, …) on a table that already has a COMPLETED index?
> >> > >> The state machine only models `ABSENT β†’ IN_PROGRESS β†’ COMPLETED`,
> >> with
> >> > >> no "config changed, rebuild" transition. We need an explicit answer
> >> > >> here β€” silently keep the old index, force a rebuild, or reject the
> >> > >> property change at DDL time. Each option has different operational
> >> > >> implications and the FIP should commit to one.
> >> > >>
> >> > >> Looking forward to your thoughts.
> >> > >>
> >> > >> Best,
> >> > >> Jark
> >> > >>
> >> > >> On Fri, 29 May 2026 at 22:41, Sagar <[email protected]>
> >> wrote:
> >> > >> >
> >> > >> > Hi ,
> >> > >> >
> >> > >> > Bumping this thread. Please take a look.
> >> > >> >
> >> > >> > Sagar.
> >> > >> >
> >> > >> > On Sat, 23 May 2026 at 9:53β€―AM, Sagar <[email protected]
> >
> >> > >> wrote:
> >> > >> >
> >> > >> > > Hi,
> >> > >> > >
> >> > >> > > I created FIP-44
> >> > >> > > <
> >> > >>
> >> >
> >>
> https://cwiki.apache.org/confluence/pages/viewpage.action?pageId=429064608
> >> > >
> >> > >> to
> >> > >> > > enhance the LanceDB integration with Fluss.
> >> > >> > >
> >> > >> > > Please review.
> >> > >> > >
> >> > >> > > Sagar.
> >> > >> > >
> >> > >>
> >> > >
> >> >
> >>
> >
>

Reply via email to