chrevanthreddy opened a new issue, #20106: URL: https://github.com/apache/hudi/issues/20106
## Summary RFC-109 vector search currently has a partitioned record-level-index (RLI) integration gap: the Spark planner always passes `false` to the vector finalist arbiter's `partitionedRecordIndex` parameter. Please help define and review a reliable way to identify partitioned RLI at query time and validate its lifecycle with vector-index reads. ## Repository facts (current RFC-109 refactor branch) - `IvfRaBitQMdtSearchAlgorithm.isPartitionedRecordIndex` currently returns `false` (`hudi-spark-common/.../analysis/HoodieVectorSearchPlanBuilder.scala`). Both approximate search and exact-rerank pass this value into `VectorIndexMdtSearchUtils` arbitration. - The existing `VectorIndexRliArbitrator` already implements partition-aware finalist grouping and passes each data-table partition to `HoodieTableMetadata.readRecordIndexLocationsWithKeys(keys, Option(partition))`. We should reuse this behavior, not implement another join/lookup. - `HoodieTableMetadata.readRecordIndexLocationsWithKeys` documents that partitioned RLI keys are not globally unique and need a data-table partition hint. The clean refactor branch has no `HoodieTableMetadata.isRecordIndexPartitioned()` API; a previous work-in-progress branch had one, but it was explicitly excluded from the initial baseline pending lifecycle verification. - This is an **integration/coverage gap**, not evidence that the existing partitioned lookup path itself is broken. For non-partitioned RLI, the existing behavior remains the baseline. ## Why this matters With partitioned RLI, looking up finalists by record key alone may be ambiguous (the same record key can occur in multiple data-table partitions), or miss the intended shard. This affects SERVE/STALE/DELETED classification and ultimately approximate exclusions and exact fetches. Until validated, we should not claim correctness for partitioned-RLI vector search. ## Proposed scope / acceptance criteria 1. Agree on a reliable source of truth for whether the *active* RLI partition is partitioned, including bootstrap, catch-up/rebuild, metadata reload, and older table compatibility. Prefer existing table/MDT config or metadata seams; avoid a new lookup protocol. 2. Wire that signal to **both** approximate and exact finalist arbitration, reusing `VectorIndexRliArbitrator`'s partition-grouped lookup and existing candidate decisions. 3. Cover identical record keys in two data partitions, moved/updated and deleted records, STALE delta vs packed candidates, a missing RLI partition, and single-slice/list-backed lookup. Verify exact positional/key fallback and approximate no-stale/no-deleted behavior. 4. Validate COW and MOR lifecycle transitions (bootstrap, writes, replay/rebuild as available), and ensure non-partitioned RLI and existing search paths remain unchanged. Report any unsupported lifecycle explicitly rather than silently treating a partitioned index as global. Related: RFC-109 vector-search proposal #19309; existing partitioned-RLI hint discussion #17116; vector delta/locator maintenance #19501. This issue is about vector finalist **routing and lifecycle detection**, not a request to replace existing RLI arbitration or the separate Tier-1 maintenance scope. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
