chrevanthreddy opened a new pull request, #20107: URL: https://github.com/apache/hudi/pull/20107
## Summary Port the proven RFC-109 IVF/RaBitQ vector-index bootstrap and Spark search implementation onto a clean current-master branch, then mechanically separate the posting scanner, RLI arbitration, and Spark read-path components. This draft is a **review/benchmark staging PR**, not a claim of completed 10M qualification. Related: RFC-109 #19309; prior incremental draft #19802 (historical implementation; please review this refactor against it rather than treating both as separate features). Partitioned-RLI integration/coverage is tracked separately in #20106. ## Included - Vector MDT Avro/payload, RaBitQ factor contract, bootstrap posting writer, generation manifest/cache, Spark SQL CREATE INDEX wiring, and option validation. - Historical packed posting scanner, delta overlay/tombstones, RaBitQ scoring, finalist RLI arbitration, and Spark approximate/exact search with positional Parquet/key-fallback and log-resident fetch paths. - Brute-force TVF filter/max-distance backward compatibility, with IVF runtime tuning options. - Mechanical read-path extractions into `VectorSearchAlgorithms`, `VectorExactSearchPlanner`, `VectorIndexRliArbitrator`, posting serde and candidate helpers. The preserved read-path was ported first, then split. ## Scope boundaries / review warnings - **Not yet 10M benchmarked on this branch.** Historical BigANN SparkApplication YAML targets Spark 4.0 / Scala 2.13 and shared writable GCS paths; this branch builds Spark 3.5 / Scala 2.12. Do not apply old manifests unchanged or attribute historical benchmark results to this branch. Current workstation GCS access requires interactive reauthentication; 10M prerequisites and isolated outputs remain unverified. - Partitioned RLI detection is intentionally hardcoded `false` in `IvfRaBitQMdtSearchAlgorithm.isPartitionedRecordIndex` for this initial baseline. The existing partition-aware arbiter path exists, but must not be claimed correct for partitioned-RLI tables until #20106 is addressed. - HFile cache telemetry/performance experiment, range prefetch, and unrelated generic core fixes are excluded. Existing defaults and SI/RLI/bootstrap behavior should not be changed by optional caching work. - This is 18 commits / 87 files from current merge base because it includes schema/DDL/bootstrap foundations as well as read path. Earlier RFC-109 PRs overlap; reviewers may prefer splitting/rebasing after benchmark validation. ## Validation done (Spark 3.5 / Scala 2.12) - Reactor `test-compile` passed (16 modules; RAT, Checkstyle, Scalastyle). - Focused Spark JUnit: planner 8/8, Parquet locator 1/1, TVF argument compatibility 5/5 (14/14 total). - Existing execution-level TVF tests: single- and batch-query combined filter/max-distance and invalid filter handling: 3/3 passed. Maven continued into unrelated ScalaTest suites; it was stopped after selected tests passed, so this is **not** a full test-suite pass. - `TestVectorIndexOptions`: 10/10; posting/RaBitQ/arbiter focused tests passed in earlier refactor slices. - Spark 3.5/Scala 2.12 shaded bundle packaged successfully at `packaging/hudi-spark-bundle/target/hudi-spark3.5-bundle_2.12-1.3.0-SNAPSHOT.jar` (local artifact; not deployed). ## Remaining review gates - Arrange GCS reauthentication and compatible Spark 3.5/Scala 2.12 cluster runtime; stage the bundle and adapted BigANN app under an isolated, immutable run ID; require source `_SUCCESS` markers. - Run 10M COW bootstrap/index/query with exact and approximate modes, collect latency/recall and inspect Spark logs plus metrics `_SUCCESS` before claiming success. Add real numbers and artifacts here. - Confirm non-partitioned RLI lifecycle and broaden regression coverage; #20106 separately tracks partitioned RLI. Please keep this PR **draft** until the 10M gate and reviewer feedback are resolved. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
