yihua opened a new issue, #20114: URL: https://github.com/apache/hudi/issues/20114
Since HUDI-9451 (#13351) the Spark file index returns one `PartitionDirectory` per file slice instead of one per partition. `HoodieFileIndex.prepareFileSlices` maps each slice to its own directory, on the snapshot path (`HoodieFileIndex.listFiles`) and the incremental path (`HoodieIncrementalFileIndex.listFiles`), for both `shouldEmbedFileSlices` modes. HUDI-9451 did that to stop shipping the whole partition's slice mapping to every task, which for partitions with tens of thousands of slices reached 100 MB+ per task. But the mapping is only attached to slices that have log files or a bootstrap base. Base-file-only slices, which is every slice of a copy-on-write table, carry plain partition values, so splitting them per slice gains nothing and changes what every consumer of `FileIndex.listFiles` sees: - `FileSourceScanExec` reports the slice count as `numPartitions` in the SQL UI and event logs. - Dynamic partition pruning and `selectedPartitions` walk one entry per slice. - Listeners and catalog stats consumers that enumerate `HadoopFsRelation.location` partitions get one entry per slice. On a table with 2,334 partitions and 15K slices a listener that logs input partitions writes 15K entries per query, and a 3.5 h ETL job over that table produced a 21 MB log line, which slowed the driver's task scheduling for the rest of the job. Expected: one directory per partition for base-file-only slices, and one directory per slice only where the slice mapping is needed. part of #20064 -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
