linliu-code opened a new pull request, #20054:
URL: https://github.com/apache/hudi/pull/20054

   ### Describe the issue this Pull Request addresses
   
   issue: #20053
   
   Token search over a text column (`WHERE message has 'disk' and 'quota'`) 
cannot use any Hudi index today, so it reads every data file. This draft is a 
copy-on-write prototype of a full-text index in the metadata table, meant to 
inform an RFC; it is not intended to merge as is.
   
   ### Summary and Changelog
   
   Adds `CREATE INDEX idx ON t USING full_text(col)` for COW tables and three 
SQL predicates (`hudi_has_token`, `hudi_has_all_tokens`, 
`hudi_has_any_tokens`). The file index uses the new index to skip base files 
that hold no row containing every query token.
   
   - New MDT partition type `FULL_TEXT_INDEX` and a nullable 
`FullTextIndexMetadata` field in `HoodieMetadata.avsc`. Per (term, base file) 
it stores a presence entry and, unless the term is dense in the file, a roaring 
bitmap of row positions.
   - `FullTextIndexer` indexes each base file a commit writes and removes the 
entries of base files it replaces (COW rewrites, clustering, insert overwrite).
   - `FullTextIndexSupport` in `HoodieFileIndex` prunes a slice only when a 
per-file coverage marker for its exact base file is visible in the index view 
being read, intersects bitmaps per file for multi-token queries, keeps slices 
with log files, and is not used for incremental or older time-travel reads. The 
marker makes pruning safe when the reader's file slices and the metadata table 
view are at different instants.
   - Tokenizer and key layout in `FullTextIndexUtils`; DDL wiring in 
`CreateIndexCommand`, `HoodieSparkIndexClient`, `HoodieIndexUtils` and index 
scheduling.
   
   ### Impact
   
   New public SQL syntax and functions, and a new metadata table partition and 
record field (additive; existing partitions are unchanged). Measured one commit 
before the coverage-marker change, on MS MARCO (8.84M passages in 16 COW base 
files, 200 real queries), queries read 2.08 of 16 files on average with 
identical results to a full scan, and median latency dropped from 6.41 s to 
0.78 s. The index is 0.68 GB. The cost is on writes: a 1% upsert that rewrites 
every file took 216 s with the index versus 22.5 s without.
   
   ### Risk Level
   
   medium. It adds a metadata table record type and field; nothing changes for 
tables without the index. Verified with `TestFullTextIndex` (co-occurrence 
fixture, COW upsert with and without entry removal, a reader and index view at 
different instants, a column added by schema evolution, time travel, 
clustering, insert overwrite, rollback, CUSTOM merge mode rejection, keyless 
tables, and a randomized test that no file with a match is ever pruned), the 
existing `TestSecondaryIndex` and `TestHoodieMetadataPayload`, and mutation 
checks on the tokenizer, the bitmap intersection, the base-instant check and 
the coverage check.
   
   Known limitations of this prototype: COW only; tokenization follows the 
JVM's Unicode tables, so writers and readers on different JDK versions could 
disagree on rare characters; renaming or dropping an indexed column is not 
blocked (base files written before the column existed, or under its old name, 
are simply never pruned); when a query has a full-text predicate, the other 
index supports are not consulted for its remaining predicates.
   
   ### Documentation Update
   
   none for this draft. A merged version would need the index documented on the 
website and an RFC.
   
   ### Contributor's checklist
   
   - [ ] Read through [contributor's 
guide](https://hudi.apache.org/contribute/how-to-contribute)
   - [ ] Enough context is provided in the sections above
   - [ ] Adequate tests were added if applicable
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to