Hi ,

If someone got a chance to look at the poc, Please let me know if this is
at a stage when I can create an fip for this.

Sagar.

On Sat, 15 Aug 2026 at 9:54 PM, Sagar <[email protected]> wrote:

> Thanks Mehul!
>
> I Wanted to share an update on the VECTOR data type for Fluss. This PR has
> an e2e POC for the same:
>
> https://github.com/apache/fluss/pull/4004.
>
> It mainly contains an IT for Lance Vector Tiering
> (LanceVectorTieringITCase) with 2 tests: testVectorTiering and
> testNullVectorTiering. I am seeing some build errors on fluss-spark but the
> IT seems to work locally.  At a high level, it:
>
>    - Writes records with vector embeddings (VECTOR(dim)) into a Fluss
>    table.
>    - A Flink background job reads those vector records from Fluss and
>    writes them into Lance storage.
>    - The test opens the resulting Lance dataset directly and checks that
>    every vector, float value, row count, and null field made it through
>    accurately without corruption or data loss.
>
> While this seems to work, there are a couple of things worth calling out:
>
> 1) Flink SQL Limitation: Flink SQL doesn't have a native VECTOR data type
> (it treats embeddings as ARRAY<FLOAT>), so vector columns aren't exposed as
> a distinct VECTOR type in SQL queries. Under the hood, Fluss models this as
> a dedicated VECTOR(dim) type mapped directly to Arrow's
> FixedSizeListVector<Float32>. This avoids the extra overhead of
> variable-sized lists and aligns 1:1 with Lance's native vector storage
> format.
>
> 2) Multi-Dimensional Array Limitation
> Currently, we only support 1D fixed-size float vectors. Multi-dimensional
> arrays (like matrices or nested arrays ARRAY<ARRAY<FLOAT>>) are not
> supported for vector tiering yet.
>
> To support that, we would need more work across schema conversion to Lance
> multi-level list format and even more things on Arrow converters. We can
> take a look at it later i think.
>
> Please take a look and see if it is at a point where we can introduce a
> FIP. I think it would be similar to this doc:
>
>
> https://docs.google.com/document/d/1idmsgMLjYScgYj-l_ABD7rwvNbDvoaq8bbMgWcgZW20/edit?tab=t.0
>
> Sagar.
>
>
> On Sat, Jul 4, 2026 at 5:45 PM Mehul Batra <[email protected]>
> wrote:
>
>> Hi Sagar,
>>
>> Thanks for the detailed findings and document, really appreciate the work
>> here.
>>
>> I think this is the right direction  as discussed on the community call,
>> let’s keep phase one focused on just the vector data type. Excited to see
>> the POC take shape, and a FIP to discuss the findings sounds like a great
>> next step.
>>
>> A few things to keep in mind from our thread & discussions:
>>
>> • Fixed dimension + fixed element type at schema time (VECTOR(1536))
>>
>> • Arrow-native layout (FixedSizeList<Float32>, zero-copy with
>> Lance/Paimon)
>>
>> • Support FLOAT32, FLOAT16, INT8 for quantization from day one so we avoid
>> a breaking change later.
>>
>> Best Regards,
>>
>> Mehul Batra
>> On Sun, Jun 28, 2026 at 10:21 AM Sagar <[email protected]> wrote:
>>
>> > Hi,
>> >
>> > Following up again to see if there are any further comments or feedback.
>> >
>> > Sagar.
>> >
>> > On Tue, 16 Jun 2026 at 11:09 PM, Sagar <[email protected]>
>> wrote:
>> >
>> > > Hi Jark
>> > >
>> > > Thanks for the detailed feedback! Please find my responses:
>> > >
>> > >
>> > > *Point 1*
>> > >
>> > > *Property-to-API mapping and unexposed parameters* — Added a mapping
>> > > table to Public Interfaces covering all three API calls. For
>> parameters
>> > > Fluss doesn't expose: replace is fixed True on both create_index() and
>> > > create_fts_index() — any other value would break the idempotency the
>> > > state machine relies on. use_tantivy is fixed False (see below).
>> > num_bits,
>> > > delete_unverified, and retrain are not exposed; LanceDB defaults
>> apply.
>> > > Since we route through JNI to lance-core (discussed below) rather than
>> > the
>> > > Python client, parameter semantics are identical to the Python
>> > equivalents
>> > > shown in the example.
>> > >
>> > > *FTS path* — Fluss targets the native FTS path (use_tantivy=False),
>> now
>> > > the LanceDB upstream default. The legacy Tantivy path is not
>> supported;
>> > it
>> > > differs in both parameter surface and on-disk format. This is fixed at
>> > the
>> > > JNI layer.
>> > >
>> > > *Default divergences* — Checked against the LanceDB docs [1]: all FTS
>> > > defaults in the FIP are identical to LanceDB's defaults, so no
>> rationale
>> > is
>> > > needed there. The only divergences are on the vector side:
>> > ef_construction
>> > > (Fluss: 150, LanceDB: 300) and lance.index.m=(none), which maps to
>> > > LanceDB's hardcoded 20. Both are now documented with rationale.
>> > >
>> > > *Side-by-side example* — Added below the existing full-configuration
>> SQL
>> > > block.
>> > >
>> > > *Point 2 — Execution model, LanceDB embedding, and horizontal scaling*
>> > >
>> > > *Where does the committer run?*
>> > >
>> > > Inside the Tiering Service worker, consistent with FIP-5. Modified the
>> > FIP
>> > > to update this
>> > >
>> > > *How is LanceDB embedded in the JVM?*
>> > >
>> > > com.lancedb:lance-core is a first-party JNI binding from the
>> > > lance-format/lance monorepo — already a dependency in Fluss's existing
>> > > Lance integration from FIP-5, not a new one introduced here.
>> > >
>> > > I verified the published 0.39.0 JAR directly. createIndex and
>> listIndexes
>> > > are present. The one gap is optimizeIndices — needed to fold newly
>> > > written rows into existing indices after each tiering cycle. The JNI
>> > > pattern is established in the codebase by nativeCreateIndex;
>> contributing
>> > > nativeOptimizeIndices is a single function addition in
>> > > java/lance-jni/src/dataset.rs with a corresponding method pair in
>> > > Dataset.java. This is a committed prerequisite of the FIP-44
>> > > implementation. No sidecar, no subprocess. FIP is updated with this
>> > detail.
>> > >
>> > > Similarly, the APIs to add an FTS index also seem missing in the jni
>> > > binding. We will need to add those as well.
>> > >
>> > > *How is index work distributed?*
>> > >
>> > > Per-table, scoped to the committer owning that table. Horizontal
>> scaling
>> > > is at table granularity.
>> > >
>> > > The single-table pinning concern is real but bounded: createIndex is
>> > > non-blocking — the build runs async inside LanceDB's Tokio runtime,
>> the
>> > > committer thread is released immediately and polls listIndexes() on
>> > > subsequent timer fires. The build itself is internally multi-threaded.
>> > The
>> > > constraint is cross-JVM-process parallelism, not single-threading. A
>> > *Scaling
>> > > Constraints* note will be added to the FIP. Coordinator-assigned index
>> > > builds are a reasonable future extension but out of scope here. This
>> is
>> > > also added to the FIP.
>> > >
>> > >
>> > > *3. Configuration drift after the table exists*
>> > >
>> > > For the initial scope of FIP-44, we will take the *'reject at DDL
>> time'*
>> > > approach.
>> > >
>> > > Index configurations will be treated as immutable once the index state
>> > > enters IN_PROGRESS or COMPLETED. If a user attempts to modify
>> properties
>> > > like lance.index.type, metric, or num_partitions via ALTER TABLE, the
>> DDL
>> > > validator will reject it.
>> > >
>> > > *Rationale:* This keeps the FIP-44 state machine strictly linear
>> (ABSENT
>> > > → IN_PROGRESS → COMPLETED). It avoids the complexities of modeling
>> > > PENDING_REBUILD states and protects Tiering workers from accidentally
>> > > triggering massive background rebuilds due to a simple property tweak.
>> > > Declarative background rebuilds for config drift can be tackled in a
>> > future
>> > > FIP. I will update the document to explicitly state this constraint
>> > >
>> > > Let me know what you think!
>> > >
>> > >
>> > > Sagar.
>> > >
>> > > [1]: https://docs.lancedb.com/search/full-text-search#advanced-usage
>> > >
>> > > On Sun, May 31, 2026 at 10:34 AM Jark Wu <[email protected]> wrote:
>> > >
>> > >> Hi Sagar,
>> > >>
>> > >> Thanks for the detailed FIP. Three comments below.
>> > >>
>> > >> ## 1. Public-interface docs need a mapping and a worked example
>> > >>
>> > >> The `lance.*` properties currently stand alone in the FIP — to
>> > >> understand any of them, a reader has to cross-reference the LanceDB
>> > >> docs. I'd like the FIP to add three things to the public-interface
>> > >> section:
>> > >>
>> > >> - An explicit table mapping each Fluss property to the LanceDB API
>> > >> call and parameter it maps to (e.g. `lance.index.type` →
>> > >> `Table.create_index(index_type=...)`).
>> > >> - An explicit mapping of each Fluss default to the corresponding
>> > >> LanceDB default, with rationale for any deliberate divergence.
>> > >> Skimming the FIP, several `lance.fts.*` defaults look like they
>> differ
>> > >> from LanceDB upstream defaults (e.g. `stem`, `remove_stop_words`,
>> > >> `ascii_folding`), and `lance.index.m`'s `(none)` effectively means
>> > >> LanceDB's hardcoded `20`. The reasons aren't stated.
>> > >> - A side-by-side example showing the same index expressed as (a) a
>> > >> Fluss `CREATE TABLE ... WITH (...)` statement, and (b) the equivalent
>> > >> LanceDB Python call. That makes the abstraction concrete for both
>> > >> reviewers and future users.
>> > >>
>> > >> Two specific things worth pinning down while you're in there:
>> > >>
>> > >> - `create_fts_index` has a legacy Tantivy path and a newer native FTS
>> > >> path (`use_tantivy=False`, now the upstream default). Which one is
>> the
>> > >> FIP targeting? The parameter surface and on-disk format both differ.
>> > >> - `Table.create_index` and `Table.optimize` have additional
>> parameters
>> > >> (`replace`, `num_bits`, `delete_unverified`, `retrain`, …) that
>> aren't
>> > >> currently mapped. Either include them or explain why they're
>> > >> deliberately hidden — `replace` in particular matters because the
>> > >> state machine relies on `create_index` being idempotent, which is
>> only
>> > >> true with `replace=True`.
>> > >>
>> > >> ## 2. Who builds the index? Execution model and horizontal scaling
>> > >>
>> > >> The FIP assigns the index lifecycle to the `LanceLakeCommitter`, but
>> > >> the deeper execution-model question is not yet answered:
>> > >>
>> > >> - Where does the committer (and therefore `create_index()` /
>> > >> `optimize()`) physically run? My reading of FIP-5 is that the
>> > >> committer lives inside the Tiering Service workers. Is that the
>> intent
>> > >> here?
>> > >>
>> > >> - If so, the Tiering Service now has to **embed LanceDB**. LanceDB is
>> > >> a Rust core with Python and Node bindings — there is no first-party
>> > >> Java client today. How is it embedded into the JVM-based tiering
>> > >> worker? JNI over the Rust core? A sidecar subprocess? Something else?
>> > >> This is a non-trivial dependency to take on and deserves explicit
>> > >> discussion in the FIP.
>> > >>
>> > >> - How is index work **distributed** across Tiering Service workers?
>> > >> Per-table affinity? Coordinator-assigned? With a single large table
>> > >> whose one heavy index takes hours to build, does the work pin to one
>> > >> worker, or can it be split? If the asynchronous build effectively
>> runs
>> > >> in-process inside the worker that initiated it, then horizontal
>> > >> scaling is per-table at best.
>> > >>
>> > >>
>> > >> ## 3. Configuration drift after the table exists
>> > >>
>> > >> What happens if a user changes `lance.index.type` (or `metric`,
>> > >> `num_partitions`, …) on a table that already has a COMPLETED index?
>> > >> The state machine only models `ABSENT → IN_PROGRESS → COMPLETED`,
>> with
>> > >> no "config changed, rebuild" transition. We need an explicit answer
>> > >> here — silently keep the old index, force a rebuild, or reject the
>> > >> property change at DDL time. Each option has different operational
>> > >> implications and the FIP should commit to one.
>> > >>
>> > >> Looking forward to your thoughts.
>> > >>
>> > >> Best,
>> > >> Jark
>> > >>
>> > >> On Fri, 29 May 2026 at 22:41, Sagar <[email protected]>
>> wrote:
>> > >> >
>> > >> > Hi ,
>> > >> >
>> > >> > Bumping this thread. Please take a look.
>> > >> >
>> > >> > Sagar.
>> > >> >
>> > >> > On Sat, 23 May 2026 at 9:53 AM, Sagar <[email protected]>
>> > >> wrote:
>> > >> >
>> > >> > > Hi,
>> > >> > >
>> > >> > > I created FIP-44
>> > >> > > <
>> > >>
>> >
>> https://cwiki.apache.org/confluence/pages/viewpage.action?pageId=429064608
>> > >
>> > >> to
>> > >> > > enhance the LanceDB integration with Fluss.
>> > >> > >
>> > >> > > Please review.
>> > >> > >
>> > >> > > Sagar.
>> > >> > >
>> > >>
>> > >
>> >
>>
>

Reply via email to