yihua opened a new pull request, #20096:
URL: https://github.com/apache/hudi/pull/20096

   ### Describe the issue this Pull Request addresses
   
   closes #20092
   part of #20064
   
   Stacked on #20069 (`HoodieEngineContext#broadcast`), review that first.
   
   Distributed metadata table lookups in `HoodieBackedTableMetadata` (column 
stats, expression, record and secondary index) capture the whole table 
metadata, with both meta clients and the file system view, and a task can list 
the data table timeline to compute the valid instants. The record index and 
metadata bloom filter functions on the write path ship the `HoodieTable` and 
open a new metadata reader in every task. The column stats read ships every 
candidate file name in every task.
   
   ### Summary and Changelog
   
   - New serializable `MetadataPartitionReader` holding only driver-resolved 
state for one metadata table partition (MDT meta client with its timeline, file 
group reader props, latest MDT instant, valid instants, file slices), exposed 
through `HoodieTableMetadata#getPartitionReader`. Distributed lookups capture 
only this reader.
   - Record index (global and partitioned) and metadata bloom filter probing 
functions hold an engine context broadcast of the reader instead of the table. 
`PartitionIdPassthrough` is static, so the lookup partitioner no longer carries 
the index and write config. Record index keys are encoded through 
`RecordIndexRawKey` before the lookup.
   - The on-cluster column stats read broadcasts its candidate file names.
   - `SparkRDDReadClient` builds a fresh table per lookup, so a reused client 
sees later commits.
   - Tests: `TestMetadataTableLookupTasks` with a counting local file system 
that fails if a lookup task lists a timeline or reads `hoodie.properties`, plus 
task payload checks; `TestMetadataTableLookupWithPlainKryo`. The I/O and 
payload tests fail on master. Follow-up: `TaskMetaIOCountingFileSystem.scala` 
here is a near duplicate of `TaskFileAccessRecordingFileSystem` added for 
#20093; merge them once both land.
   
   ### Impact
   
   Smaller task binaries for metadata table lookups (about half on a local test 
table, more with cluster Hadoop configurations or many candidate files) and no 
timeline or table config I/O in lookup tasks. Write-path index lookups use the 
driver's metadata table snapshot for all tasks. 
`PartitionedRecordIndexFileGroupLookupFunction` and 
`HoodieMetadataBloomFilterProbingFunction` constructors change.
   
   With `hoodie.metadata.metrics.enable=true`, bloom filter probing tasks no 
longer record `lookup_meta_index_bloom_filters` and its file count from 
executor JVMs, since they no longer build a table metadata (and a metrics 
reporter) there. These values never reached the driver's registry and were only 
visible through a push reporter reachable from executors. Driver-side metadata 
metrics and the record index lookup metrics are unchanged.
   
   ### Risk Level
   
   medium. The lookup logic is moved, not rewritten; results are covered by the 
existing index suites plus new equivalence tests. Write-path lookups now read 
one driver snapshot of the metadata table. This PR conflicts textually with the 
PR for #20093 in `HoodieBackedTableMetadata#getRecordsByKeyPrefixes`; whichever 
merges second rebases.
   
   ### Documentation Update
   
   none
   
   ### Contributor's checklist
   
   - [ ] Read through [contributor's 
guide](https://hudi.apache.org/contribute/how-to-contribute)
   - [ ] Enough context is provided in the sections above
   - [ ] Adequate tests were added if applicable
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to