linliu-code opened a new pull request, #20054: URL: https://github.com/apache/hudi/pull/20054
### Describe the issue this Pull Request addresses issue: #20053 Token search over a text column (`WHERE message has 'disk' and 'quota'`) cannot use any Hudi index today, so it reads every data file. This draft is a copy-on-write prototype of a full-text index in the metadata table, meant to inform an RFC; it is not intended to merge as is. ### Summary and Changelog Adds `CREATE INDEX idx ON t USING full_text(col)` for COW tables and three SQL predicates (`hudi_has_token`, `hudi_has_all_tokens`, `hudi_has_any_tokens`). The file index uses the new index to skip base files that hold no row containing every query token. - New MDT partition type `FULL_TEXT_INDEX` and a nullable `FullTextIndexMetadata` field in `HoodieMetadata.avsc`. Per (term, base file) it stores a presence entry and, unless the term is dense in the file, a roaring bitmap of row positions. - `FullTextIndexer` indexes each base file a commit writes and removes the entries of base files it replaces (COW rewrites, clustering, insert overwrite). - `FullTextIndexSupport` in `HoodieFileIndex` prunes a slice only when a per-file coverage marker for its exact base file is visible in the index view being read, intersects bitmaps per file for multi-token queries, keeps slices with log files, and is not used for incremental or older time-travel reads. The marker makes pruning safe when the reader's file slices and the metadata table view are at different instants. - Tokenizer and key layout in `FullTextIndexUtils`; DDL wiring in `CreateIndexCommand`, `HoodieSparkIndexClient`, `HoodieIndexUtils` and index scheduling. ### Impact New public SQL syntax and functions, and a new metadata table partition and record field (additive; existing partitions are unchanged). Measured one commit before the coverage-marker change, on MS MARCO (8.84M passages in 16 COW base files, 200 real queries), queries read 2.08 of 16 files on average with identical results to a full scan, and median latency dropped from 6.41 s to 0.78 s. The index is 0.68 GB. The cost is on writes: a 1% upsert that rewrites every file took 216 s with the index versus 22.5 s without. ### Risk Level medium. It adds a metadata table record type and field; nothing changes for tables without the index. Verified with `TestFullTextIndex` (co-occurrence fixture, COW upsert with and without entry removal, a reader and index view at different instants, a column added by schema evolution, time travel, clustering, insert overwrite, rollback, CUSTOM merge mode rejection, keyless tables, and a randomized test that no file with a match is ever pruned), the existing `TestSecondaryIndex` and `TestHoodieMetadataPayload`, and mutation checks on the tokenizer, the bitmap intersection, the base-instant check and the coverage check. Known limitations of this prototype: COW only; tokenization follows the JVM's Unicode tables, so writers and readers on different JDK versions could disagree on rare characters; renaming or dropping an indexed column is not blocked (base files written before the column existed, or under its old name, are simply never pruned); when a query has a full-text predicate, the other index supports are not consulted for its remaining predicates. ### Documentation Update none for this draft. A merged version would need the index documented on the website and an RFC. ### Contributor's checklist - [ ] Read through [contributor's guide](https://hudi.apache.org/contribute/how-to-contribute) - [ ] Enough context is provided in the sections above - [ ] Adequate tests were added if applicable -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
