chrevanthreddy opened a new issue, #20106:
URL: https://github.com/apache/hudi/issues/20106

   ## Summary
   RFC-109 vector search currently has a partitioned record-level-index (RLI) 
integration gap: the Spark planner always passes `false` to the vector finalist 
arbiter's `partitionedRecordIndex` parameter. Please help define and review a 
reliable way to identify partitioned RLI at query time and validate its 
lifecycle with vector-index reads.
   
   ## Repository facts (current RFC-109 refactor branch)
   - `IvfRaBitQMdtSearchAlgorithm.isPartitionedRecordIndex` currently returns 
`false` (`hudi-spark-common/.../analysis/HoodieVectorSearchPlanBuilder.scala`). 
Both approximate search and exact-rerank pass this value into 
`VectorIndexMdtSearchUtils` arbitration.
   - The existing `VectorIndexRliArbitrator` already implements partition-aware 
finalist grouping and passes each data-table partition to 
`HoodieTableMetadata.readRecordIndexLocationsWithKeys(keys, 
Option(partition))`. We should reuse this behavior, not implement another 
join/lookup.
   - `HoodieTableMetadata.readRecordIndexLocationsWithKeys` documents that 
partitioned RLI keys are not globally unique and need a data-table partition 
hint. The clean refactor branch has no 
`HoodieTableMetadata.isRecordIndexPartitioned()` API; a previous 
work-in-progress branch had one, but it was explicitly excluded from the 
initial baseline pending lifecycle verification.
   - This is an **integration/coverage gap**, not evidence that the existing 
partitioned lookup path itself is broken. For non-partitioned RLI, the existing 
behavior remains the baseline.
   
   ## Why this matters
   With partitioned RLI, looking up finalists by record key alone may be 
ambiguous (the same record key can occur in multiple data-table partitions), or 
miss the intended shard. This affects SERVE/STALE/DELETED classification and 
ultimately approximate exclusions and exact fetches. Until validated, we should 
not claim correctness for partitioned-RLI vector search.
   
   ## Proposed scope / acceptance criteria
   1. Agree on a reliable source of truth for whether the *active* RLI 
partition is partitioned, including bootstrap, catch-up/rebuild, metadata 
reload, and older table compatibility. Prefer existing table/MDT config or 
metadata seams; avoid a new lookup protocol.
   2. Wire that signal to **both** approximate and exact finalist arbitration, 
reusing `VectorIndexRliArbitrator`'s partition-grouped lookup and existing 
candidate decisions.
   3. Cover identical record keys in two data partitions, moved/updated and 
deleted records, STALE delta vs packed candidates, a missing RLI partition, and 
single-slice/list-backed lookup. Verify exact positional/key fallback and 
approximate no-stale/no-deleted behavior.
   4. Validate COW and MOR lifecycle transitions (bootstrap, writes, 
replay/rebuild as available), and ensure non-partitioned RLI and existing 
search paths remain unchanged. Report any unsupported lifecycle explicitly 
rather than silently treating a partitioned index as global.
   
   Related: RFC-109 vector-search proposal #19309; existing partitioned-RLI 
hint discussion #17116; vector delta/locator maintenance #19501. This issue is 
about vector finalist **routing and lifecycle detection**, not a request to 
replace existing RLI arbitration or the separate Tier-1 maintenance scope.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to