yihua opened a new pull request, #20115:
URL: https://github.com/apache/hudi/pull/20115

   ### Describe the issue this Pull Request addresses
   
   closes #20114
   part of #20064
   
   Since HUDI-9451 (#13351) `HoodieFileIndex.prepareFileSlices` returns one 
`PartitionDirectory` per file slice, on the snapshot and incremental paths and 
in both `shouldEmbedFileSlices` modes. HUDI-9451 did that so a task is shipped 
only the slice it reads instead of the whole partition's slice mapping, which 
for partitions with tens of thousands of slices reached 100 MB+ per task. The 
mapping is only attached to slices with log files or a bootstrap base, though. 
Base-file-only slices, which is every slice of a copy-on-write table, carry 
plain partition values, so splitting them per slice gains nothing and changes 
what every consumer of `FileIndex.listFiles` sees: `FileSourceScanExec` reports 
the slice count as `numPartitions`, dynamic partition pruning walks one entry 
per slice, and listeners or stats consumers that enumerate the relation's 
partitions get one entry per slice. On a table with 2,334 partitions and 15K 
slices a query listener that logs input partitions wrote 15K entri
 es per query, and over a long ETL job a 21 MB line that slowed the driver's 
task scheduling.
   
   ### Summary and Changelog
   
   `PartitionDirectoryConverter.convertFileSlicesToPartitionDirectories` takes 
the slices of one partition and returns one directory with plain partition 
values holding the delegate files of all base-file-only slices, plus one 
directory per slice that has log files or a bootstrap base, each with a 
single-entry `HoodiePartitionFileSliceMapping` as before. `prepareFileSlices` 
takes the slices grouped by partition values, which is how both `listFiles` 
callers already produce them, and the non-embedded branch puts all files of a 
partition into one directory again. Non-partitioned tables get the same 
grouping under their single empty partition value.
   
   Split planning is unchanged: Spark flattens the directories to files, sorts 
them by length and bin-packs them, and the file order is preserved for 
base-file-only slices. On a mixed partition the base-file-only directory comes 
before the per-slice directories, so files of equal length may swap places 
between tasks, with the same task count and task sizes.
   
   Tests: 
`TestPartitionDirectoryConverter.testConvertFileSlicesToPartitionDirectories` 
checks the directory shapes for a partition mixing base-file-only, base plus 
log, log-only and bootstrap slices. 
`TestHoodieFileIndex.testListFilesGroupsFileSlicesByPartition` writes 3 file 
groups into each of 5 partitions of a COW and a MOR table and checks that 
`listFiles` returns 5 directories, that after an update the two file groups 
with log files get their own directories, and that the file partitions Spark 
plans match those planned from one directory per slice. Without the fix it 
fails with `expected: <5> but was: <15>`.
   
   ### Impact
   
   `listFiles` returns one directory per partition for copy-on-write tables and 
for the base-file-only slices of merge-on-read tables. `numPartitions` in the 
scan metrics reports partitions again. Fewer driver objects per scan.
   
   ### Risk Level
   
   low. Slices that need the mapping keep exactly the per-slice shape of 
HUDI-9451, and the executor read path is untouched.
   
   ### Documentation Update
   
   none
   
   ### Contributor's checklist
   
   - [ ] Read through [contributor's 
guide](https://hudi.apache.org/contribute/how-to-contribute)
   - [ ] Enough context is provided in the sections above
   - [ ] Adequate tests were added if applicable
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to