Hi Jark Thanks for the detailed feedback! Please find my responses:
*Point 1* *Property-to-API mapping and unexposed parameters* — Added a mapping table to Public Interfaces covering all three API calls. For parameters Fluss doesn't expose: replace is fixed True on both create_index() and create_fts_index() — any other value would break the idempotency the state machine relies on. use_tantivy is fixed False (see below). num_bits, delete_unverified, and retrain are not exposed; LanceDB defaults apply. Since we route through JNI to lance-core (discussed below) rather than the Python client, parameter semantics are identical to the Python equivalents shown in the example. *FTS path* — Fluss targets the native FTS path (use_tantivy=False), now the LanceDB upstream default. The legacy Tantivy path is not supported; it differs in both parameter surface and on-disk format. This is fixed at the JNI layer. *Default divergences* — Checked against the LanceDB docs [1]: all FTS defaults in the FIP are identical to LanceDB's defaults, so no rationale is needed there. The only divergences are on the vector side: ef_construction (Fluss: 150, LanceDB: 300) and lance.index.m=(none), which maps to LanceDB's hardcoded 20. Both are now documented with rationale. *Side-by-side example* — Added below the existing full-configuration SQL block. *Point 2 — Execution model, LanceDB embedding, and horizontal scaling* *Where does the committer run?* Inside the Tiering Service worker, consistent with FIP-5. Modified the FIP to update this *How is LanceDB embedded in the JVM?* com.lancedb:lance-core is a first-party JNI binding from the lance-format/lance monorepo — already a dependency in Fluss's existing Lance integration from FIP-5, not a new one introduced here. I verified the published 0.39.0 JAR directly. createIndex and listIndexes are present. The one gap is optimizeIndices — needed to fold newly written rows into existing indices after each tiering cycle. The JNI pattern is established in the codebase by nativeCreateIndex; contributing nativeOptimizeIndices is a single function addition in java/lance-jni/src/dataset.rs with a corresponding method pair in Dataset.java. This is a committed prerequisite of the FIP-44 implementation. No sidecar, no subprocess. FIP is updated with this detail. Similarly, the APIs to add an FTS index also seem missing in the jni binding. We will need to add those as well. *How is index work distributed?* Per-table, scoped to the committer owning that table. Horizontal scaling is at table granularity. The single-table pinning concern is real but bounded: createIndex is non-blocking — the build runs async inside LanceDB's Tokio runtime, the committer thread is released immediately and polls listIndexes() on subsequent timer fires. The build itself is internally multi-threaded. The constraint is cross-JVM-process parallelism, not single-threading. A *Scaling Constraints* note will be added to the FIP. Coordinator-assigned index builds are a reasonable future extension but out of scope here. This is also added to the FIP. *3. Configuration drift after the table exists* For the initial scope of FIP-44, we will take the *'reject at DDL time'* approach. Index configurations will be treated as immutable once the index state enters IN_PROGRESS or COMPLETED. If a user attempts to modify properties like lance.index.type, metric, or num_partitions via ALTER TABLE, the DDL validator will reject it. *Rationale:* This keeps the FIP-44 state machine strictly linear (ABSENT → IN_PROGRESS → COMPLETED). It avoids the complexities of modeling PENDING_REBUILD states and protects Tiering workers from accidentally triggering massive background rebuilds due to a simple property tweak. Declarative background rebuilds for config drift can be tackled in a future FIP. I will update the document to explicitly state this constraint Let me know what you think! Sagar. [1]: https://docs.lancedb.com/search/full-text-search#advanced-usage On Sun, May 31, 2026 at 10:34 AM Jark Wu <[email protected]> wrote: > Hi Sagar, > > Thanks for the detailed FIP. Three comments below. > > ## 1. Public-interface docs need a mapping and a worked example > > The `lance.*` properties currently stand alone in the FIP — to > understand any of them, a reader has to cross-reference the LanceDB > docs. I'd like the FIP to add three things to the public-interface > section: > > - An explicit table mapping each Fluss property to the LanceDB API > call and parameter it maps to (e.g. `lance.index.type` → > `Table.create_index(index_type=...)`). > - An explicit mapping of each Fluss default to the corresponding > LanceDB default, with rationale for any deliberate divergence. > Skimming the FIP, several `lance.fts.*` defaults look like they differ > from LanceDB upstream defaults (e.g. `stem`, `remove_stop_words`, > `ascii_folding`), and `lance.index.m`'s `(none)` effectively means > LanceDB's hardcoded `20`. The reasons aren't stated. > - A side-by-side example showing the same index expressed as (a) a > Fluss `CREATE TABLE ... WITH (...)` statement, and (b) the equivalent > LanceDB Python call. That makes the abstraction concrete for both > reviewers and future users. > > Two specific things worth pinning down while you're in there: > > - `create_fts_index` has a legacy Tantivy path and a newer native FTS > path (`use_tantivy=False`, now the upstream default). Which one is the > FIP targeting? The parameter surface and on-disk format both differ. > - `Table.create_index` and `Table.optimize` have additional parameters > (`replace`, `num_bits`, `delete_unverified`, `retrain`, …) that aren't > currently mapped. Either include them or explain why they're > deliberately hidden — `replace` in particular matters because the > state machine relies on `create_index` being idempotent, which is only > true with `replace=True`. > > ## 2. Who builds the index? Execution model and horizontal scaling > > The FIP assigns the index lifecycle to the `LanceLakeCommitter`, but > the deeper execution-model question is not yet answered: > > - Where does the committer (and therefore `create_index()` / > `optimize()`) physically run? My reading of FIP-5 is that the > committer lives inside the Tiering Service workers. Is that the intent > here? > > - If so, the Tiering Service now has to **embed LanceDB**. LanceDB is > a Rust core with Python and Node bindings — there is no first-party > Java client today. How is it embedded into the JVM-based tiering > worker? JNI over the Rust core? A sidecar subprocess? Something else? > This is a non-trivial dependency to take on and deserves explicit > discussion in the FIP. > > - How is index work **distributed** across Tiering Service workers? > Per-table affinity? Coordinator-assigned? With a single large table > whose one heavy index takes hours to build, does the work pin to one > worker, or can it be split? If the asynchronous build effectively runs > in-process inside the worker that initiated it, then horizontal > scaling is per-table at best. > > > ## 3. Configuration drift after the table exists > > What happens if a user changes `lance.index.type` (or `metric`, > `num_partitions`, …) on a table that already has a COMPLETED index? > The state machine only models `ABSENT → IN_PROGRESS → COMPLETED`, with > no "config changed, rebuild" transition. We need an explicit answer > here — silently keep the old index, force a rebuild, or reject the > property change at DDL time. Each option has different operational > implications and the FIP should commit to one. > > Looking forward to your thoughts. > > Best, > Jark > > On Fri, 29 May 2026 at 22:41, Sagar <[email protected]> wrote: > > > > Hi , > > > > Bumping this thread. Please take a look. > > > > Sagar. > > > > On Sat, 23 May 2026 at 9:53 AM, Sagar <[email protected]> wrote: > > > > > Hi, > > > > > > I created FIP-44 > > > < > https://cwiki.apache.org/confluence/pages/viewpage.action?pageId=429064608> > to > > > enhance the LanceDB integration with Fluss. > > > > > > Please review. > > > > > > Sagar. > > > >
