qzyu999 opened a new issue, #1637:
URL: https://github.com/apache/datafusion-python/issues/1637

   ## Is your feature request related to a problem or challenge?
   
   `DataFrame.write_parquet()` currently returns `None`. After writing, there 
is no way to retrieve per-file metadata (row counts, byte sizes, column 
statistics) for the files that were produced. This forces consumers that need 
file-level statistics — such as Apache Iceberg, Delta Lake, and Apache Hudi — 
to either:
   
   1. Re-read Parquet footers from object storage after writing (extra I/O 
round-trips)
   2. Bypass DataFusion's write pipeline entirely and use PyArrow's 
`ParquetWriter` with `metadata_collector`
   
   This is a blocker for building a complete DataFusion-based write backend for 
table formats that require per-file column statistics in their commit metadata 
(e.g., Iceberg's `DataFile` entries need `column_sizes`, `null_counts`, 
`lower_bounds`, `upper_bounds`, `split_offsets`).
   
   ## Describe the solution you'd like
   
   After 
[apache/datafusion#23472](https://github.com/apache/datafusion/issues/23472) / 
[apache/datafusion#23656](https://github.com/apache/datafusion/pull/23656) 
lands in the Rust core, `ParquetSink` will expose a `file_metadata()` method 
returning per-file path, row count, and byte size. The Python bindings should 
surface this:
   
   ```python
   # Option A: write_parquet returns metadata directly
   metadata = df.write_parquet("/path/to/output/")
   # metadata: list[dict] = [
   #     {"path": "part-0.parquet", "row_count": 500, "byte_size": 4096},
   #     {"path": "part-1.parquet", "row_count": 500, "byte_size": 3840},
   # ]
   
   # Option B: write_parquet returns a WriteResult object
   result = df.write_parquet("/path/to/output/")
   result.count        # 1000
   result.file_metadata  # list of per-file metadata dicts
   ```
   
   At minimum, each file metadata entry should include:
   - `path` (str): Object-store path of the written file
   - `row_count` (int): Number of rows in this file
   - `byte_size` (int): Sum of compressed row group sizes
   
   Optionally (for full table-format integration):
   - `metadata` (bytes | None): Serialized Parquet `FileMetaData` (Thrift 
compact), enabling consumers to extract column statistics without re-reading 
the file
   
   ## Describe alternatives you've considered
   
   - **Return just the count** (status quo): Insufficient for table format 
integration.
   - **Expose via a separate accessor**: e.g. `ctx.last_write_metadata()` — 
awkward API, not composable.
   - **Return raw bytes of the full Parquet footer**: Maximally informative but 
heavier. A structured dict with optional raw bytes is more ergonomic.
   
   ## Additional context
   
   - **Upstream dependency:** 
[apache/datafusion#23656](https://github.com/apache/datafusion/pull/23656) adds 
`DataSink::file_metadata()` to the Rust core. This issue tracks exposing it 
through the Python bindings.
   - **Motivation:** PyIceberg is building a [pluggable execution 
backend](https://github.com/apache/iceberg-python) with DataFusion for 
bounded-memory operations. A DataFusion write backend would enable single-pass 
Copy-on-Write deletes (read → filter → write entirely in Rust with 
spill-to-disk), but requires per-file metadata to construct Iceberg `DataFile` 
commit entries.
   - **Related:** #1624 (per-session object store config) is the other piece 
needed for a complete DataFusion write backend in PyIceberg.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to