linliu-code opened a new issue, #20053: URL: https://github.com/apache/hudi/issues/20053
### Feature Description **What the feature achieves:** Queries that filter on words in a text column (for example, find rows whose `message` contains the tokens `disk` and `quota`) cannot use any Hudi index today, so they read every data file in the pruned partitions. A full-text index in the metadata table would let the file index skip data files that cannot contain a match. **Why this feature is needed:** Log, event and document tables commonly carry large text columns, and token search over them is a frequent query shape. Column stats and the secondary index do not help: the predicate is on tokens inside the value, not on the whole value or a range. ### User Experience **How users will use this feature:** - Configuration changes needed: none beyond enabling the metadata table. - API changes: `CREATE INDEX idx ON t USING full_text(col) [OPTIONS (...)]`, and SQL predicates `hudi_has_token(col, token)`, `hudi_has_all_tokens(col, query)`, `hudi_has_any_tokens(col, query)`. - Usage examples: `SELECT * FROM t WHERE hudi_has_all_tokens(message, 'disk quota')`. ### Hudi RFC Requirements **RFC PR link:** none yet. **Why RFC is/isn't needed:** - Does this change public interfaces/APIs? Yes (new `USING full_text` index type and three SQL functions). - Does this change storage format? Yes (a new metadata table partition type and record field). - Justification: an RFC is needed before anything merges. The linked draft PR is a copy-on-write prototype meant to inform it: it prunes files by row-level token co-occurrence using per-file postings in the metadata table. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
