Thanks Mehul!

I Wanted to share an update on the VECTOR data type for Fluss. This PR has
an e2e POC for the same:

https://github.com/apache/fluss/pull/4004.

It mainly contains an IT for Lance Vector Tiering
(LanceVectorTieringITCase) with 2 tests: testVectorTiering and
testNullVectorTiering. I am seeing some build errors on fluss-spark but the
IT seems to work locally.  At a high level, it:

   - Writes records with vector embeddings (VECTOR(dim)) into a Fluss table.
   - A Flink background job reads those vector records from Fluss and
   writes them into Lance storage.
   - The test opens the resulting Lance dataset directly and checks that
   every vector, float value, row count, and null field made it through
   accurately without corruption or data loss.

While this seems to work, there are a couple of things worth calling out:

1) Flink SQL Limitation: Flink SQL doesn't have a native VECTOR data type
(it treats embeddings as ARRAY<FLOAT>), so vector columns aren't exposed as
a distinct VECTOR type in SQL queries. Under the hood, Fluss models this as
a dedicated VECTOR(dim) type mapped directly to Arrow's
FixedSizeListVector<Float32>. This avoids the extra overhead of
variable-sized lists and aligns 1:1 with Lance's native vector storage
format.

2) Multi-Dimensional Array Limitation
Currently, we only support 1D fixed-size float vectors. Multi-dimensional
arrays (like matrices or nested arrays ARRAY<ARRAY<FLOAT>>) are not
supported for vector tiering yet.

To support that, we would need more work across schema conversion to Lance
multi-level list format and even more things on Arrow converters. We can
take a look at it later i think.

Please take a look and see if it is at a point where we can introduce a
FIP. I think it would be similar to this doc:

https://docs.google.com/document/d/1idmsgMLjYScgYj-l_ABD7rwvNbDvoaq8bbMgWcgZW20/edit?tab=t.0

Sagar.


On Sat, Jul 4, 2026 at 5:45 PM Mehul Batra <[email protected]> wrote:

> Hi Sagar,
>
> Thanks for the detailed findings and document, really appreciate the work
> here.
>
> I think this is the right direction  as discussed on the community call,
> let’s keep phase one focused on just the vector data type. Excited to see
> the POC take shape, and a FIP to discuss the findings sounds like a great
> next step.
>
> A few things to keep in mind from our thread & discussions:
>
> • Fixed dimension + fixed element type at schema time (VECTOR(1536))
>
> • Arrow-native layout (FixedSizeList<Float32>, zero-copy with Lance/Paimon)
>
> • Support FLOAT32, FLOAT16, INT8 for quantization from day one so we avoid
> a breaking change later.
>
> Best Regards,
>
> Mehul Batra
> On Sun, Jun 28, 2026 at 10:21 AM Sagar <[email protected]> wrote:
>
> > Hi,
> >
> > Following up again to see if there are any further comments or feedback.
> >
> > Sagar.
> >
> > On Tue, 16 Jun 2026 at 11:09 PM, Sagar <[email protected]>
> wrote:
> >
> > > Hi Jark
> > >
> > > Thanks for the detailed feedback! Please find my responses:
> > >
> > >
> > > *Point 1*
> > >
> > > *Property-to-API mapping and unexposed parameters* — Added a mapping
> > > table to Public Interfaces covering all three API calls. For parameters
> > > Fluss doesn't expose: replace is fixed True on both create_index() and
> > > create_fts_index() — any other value would break the idempotency the
> > > state machine relies on. use_tantivy is fixed False (see below).
> > num_bits,
> > > delete_unverified, and retrain are not exposed; LanceDB defaults apply.
> > > Since we route through JNI to lance-core (discussed below) rather than
> > the
> > > Python client, parameter semantics are identical to the Python
> > equivalents
> > > shown in the example.
> > >
> > > *FTS path* — Fluss targets the native FTS path (use_tantivy=False), now
> > > the LanceDB upstream default. The legacy Tantivy path is not supported;
> > it
> > > differs in both parameter surface and on-disk format. This is fixed at
> > the
> > > JNI layer.
> > >
> > > *Default divergences* — Checked against the LanceDB docs [1]: all FTS
> > > defaults in the FIP are identical to LanceDB's defaults, so no
> rationale
> > is
> > > needed there. The only divergences are on the vector side:
> > ef_construction
> > > (Fluss: 150, LanceDB: 300) and lance.index.m=(none), which maps to
> > > LanceDB's hardcoded 20. Both are now documented with rationale.
> > >
> > > *Side-by-side example* — Added below the existing full-configuration
> SQL
> > > block.
> > >
> > > *Point 2 — Execution model, LanceDB embedding, and horizontal scaling*
> > >
> > > *Where does the committer run?*
> > >
> > > Inside the Tiering Service worker, consistent with FIP-5. Modified the
> > FIP
> > > to update this
> > >
> > > *How is LanceDB embedded in the JVM?*
> > >
> > > com.lancedb:lance-core is a first-party JNI binding from the
> > > lance-format/lance monorepo — already a dependency in Fluss's existing
> > > Lance integration from FIP-5, not a new one introduced here.
> > >
> > > I verified the published 0.39.0 JAR directly. createIndex and
> listIndexes
> > > are present. The one gap is optimizeIndices — needed to fold newly
> > > written rows into existing indices after each tiering cycle. The JNI
> > > pattern is established in the codebase by nativeCreateIndex;
> contributing
> > > nativeOptimizeIndices is a single function addition in
> > > java/lance-jni/src/dataset.rs with a corresponding method pair in
> > > Dataset.java. This is a committed prerequisite of the FIP-44
> > > implementation. No sidecar, no subprocess. FIP is updated with this
> > detail.
> > >
> > > Similarly, the APIs to add an FTS index also seem missing in the jni
> > > binding. We will need to add those as well.
> > >
> > > *How is index work distributed?*
> > >
> > > Per-table, scoped to the committer owning that table. Horizontal
> scaling
> > > is at table granularity.
> > >
> > > The single-table pinning concern is real but bounded: createIndex is
> > > non-blocking — the build runs async inside LanceDB's Tokio runtime, the
> > > committer thread is released immediately and polls listIndexes() on
> > > subsequent timer fires. The build itself is internally multi-threaded.
> > The
> > > constraint is cross-JVM-process parallelism, not single-threading. A
> > *Scaling
> > > Constraints* note will be added to the FIP. Coordinator-assigned index
> > > builds are a reasonable future extension but out of scope here. This is
> > > also added to the FIP.
> > >
> > >
> > > *3. Configuration drift after the table exists*
> > >
> > > For the initial scope of FIP-44, we will take the *'reject at DDL
> time'*
> > > approach.
> > >
> > > Index configurations will be treated as immutable once the index state
> > > enters IN_PROGRESS or COMPLETED. If a user attempts to modify
> properties
> > > like lance.index.type, metric, or num_partitions via ALTER TABLE, the
> DDL
> > > validator will reject it.
> > >
> > > *Rationale:* This keeps the FIP-44 state machine strictly linear
> (ABSENT
> > > → IN_PROGRESS → COMPLETED). It avoids the complexities of modeling
> > > PENDING_REBUILD states and protects Tiering workers from accidentally
> > > triggering massive background rebuilds due to a simple property tweak.
> > > Declarative background rebuilds for config drift can be tackled in a
> > future
> > > FIP. I will update the document to explicitly state this constraint
> > >
> > > Let me know what you think!
> > >
> > >
> > > Sagar.
> > >
> > > [1]: https://docs.lancedb.com/search/full-text-search#advanced-usage
> > >
> > > On Sun, May 31, 2026 at 10:34 AM Jark Wu <[email protected]> wrote:
> > >
> > >> Hi Sagar,
> > >>
> > >> Thanks for the detailed FIP. Three comments below.
> > >>
> > >> ## 1. Public-interface docs need a mapping and a worked example
> > >>
> > >> The `lance.*` properties currently stand alone in the FIP — to
> > >> understand any of them, a reader has to cross-reference the LanceDB
> > >> docs. I'd like the FIP to add three things to the public-interface
> > >> section:
> > >>
> > >> - An explicit table mapping each Fluss property to the LanceDB API
> > >> call and parameter it maps to (e.g. `lance.index.type` →
> > >> `Table.create_index(index_type=...)`).
> > >> - An explicit mapping of each Fluss default to the corresponding
> > >> LanceDB default, with rationale for any deliberate divergence.
> > >> Skimming the FIP, several `lance.fts.*` defaults look like they differ
> > >> from LanceDB upstream defaults (e.g. `stem`, `remove_stop_words`,
> > >> `ascii_folding`), and `lance.index.m`'s `(none)` effectively means
> > >> LanceDB's hardcoded `20`. The reasons aren't stated.
> > >> - A side-by-side example showing the same index expressed as (a) a
> > >> Fluss `CREATE TABLE ... WITH (...)` statement, and (b) the equivalent
> > >> LanceDB Python call. That makes the abstraction concrete for both
> > >> reviewers and future users.
> > >>
> > >> Two specific things worth pinning down while you're in there:
> > >>
> > >> - `create_fts_index` has a legacy Tantivy path and a newer native FTS
> > >> path (`use_tantivy=False`, now the upstream default). Which one is the
> > >> FIP targeting? The parameter surface and on-disk format both differ.
> > >> - `Table.create_index` and `Table.optimize` have additional parameters
> > >> (`replace`, `num_bits`, `delete_unverified`, `retrain`, …) that aren't
> > >> currently mapped. Either include them or explain why they're
> > >> deliberately hidden — `replace` in particular matters because the
> > >> state machine relies on `create_index` being idempotent, which is only
> > >> true with `replace=True`.
> > >>
> > >> ## 2. Who builds the index? Execution model and horizontal scaling
> > >>
> > >> The FIP assigns the index lifecycle to the `LanceLakeCommitter`, but
> > >> the deeper execution-model question is not yet answered:
> > >>
> > >> - Where does the committer (and therefore `create_index()` /
> > >> `optimize()`) physically run? My reading of FIP-5 is that the
> > >> committer lives inside the Tiering Service workers. Is that the intent
> > >> here?
> > >>
> > >> - If so, the Tiering Service now has to **embed LanceDB**. LanceDB is
> > >> a Rust core with Python and Node bindings — there is no first-party
> > >> Java client today. How is it embedded into the JVM-based tiering
> > >> worker? JNI over the Rust core? A sidecar subprocess? Something else?
> > >> This is a non-trivial dependency to take on and deserves explicit
> > >> discussion in the FIP.
> > >>
> > >> - How is index work **distributed** across Tiering Service workers?
> > >> Per-table affinity? Coordinator-assigned? With a single large table
> > >> whose one heavy index takes hours to build, does the work pin to one
> > >> worker, or can it be split? If the asynchronous build effectively runs
> > >> in-process inside the worker that initiated it, then horizontal
> > >> scaling is per-table at best.
> > >>
> > >>
> > >> ## 3. Configuration drift after the table exists
> > >>
> > >> What happens if a user changes `lance.index.type` (or `metric`,
> > >> `num_partitions`, …) on a table that already has a COMPLETED index?
> > >> The state machine only models `ABSENT → IN_PROGRESS → COMPLETED`, with
> > >> no "config changed, rebuild" transition. We need an explicit answer
> > >> here — silently keep the old index, force a rebuild, or reject the
> > >> property change at DDL time. Each option has different operational
> > >> implications and the FIP should commit to one.
> > >>
> > >> Looking forward to your thoughts.
> > >>
> > >> Best,
> > >> Jark
> > >>
> > >> On Fri, 29 May 2026 at 22:41, Sagar <[email protected]>
> wrote:
> > >> >
> > >> > Hi ,
> > >> >
> > >> > Bumping this thread. Please take a look.
> > >> >
> > >> > Sagar.
> > >> >
> > >> > On Sat, 23 May 2026 at 9:53 AM, Sagar <[email protected]>
> > >> wrote:
> > >> >
> > >> > > Hi,
> > >> > >
> > >> > > I created FIP-44
> > >> > > <
> > >>
> >
> https://cwiki.apache.org/confluence/pages/viewpage.action?pageId=429064608
> > >
> > >> to
> > >> > > enhance the LanceDB integration with Fluss.
> > >> > >
> > >> > > Please review.
> > >> > >
> > >> > > Sagar.
> > >> > >
> > >>
> > >
> >
>

Reply via email to