andygrove opened a new issue, #6146: URL: https://github.com/apache/datafusion-comet/issues/6146
### Describe the bug For float and double fields inside a list or map, the native Iceberg writer can over-count NaN values in the data file's metrics when the input batch is already a slice of a larger array. The writer's `RowSlicer` handles slices it cuts itself. A batch that arrives already sliced (for example from `GlobalLimitExec` in `INSERT ... SELECT ... LIMIT n`) keeps its full child arrays through the schema cast (`iceberg_write.rs:871-879`), and iceberg-rust's NaN counter reads `list.values()`, so it also counts NaNs outside the batch's window. Effect: wrong `nan_value_counts` for nested float/double fields. Spark cannot push predicates onto list elements, so this is mostly incorrect metadata rather than wrong query results. Found by reading the code; not reproduced. ### Steps to reproduce ```sql CREATE TABLE t (id INT, xs ARRAY<DOUBLE>) USING iceberg; -- source where NaNs appear only in rows past the limit INSERT INTO t SELECT id, xs FROM src ORDER BY id LIMIT 10; SELECT readable_metrics FROM t.files; ``` ### Expected behavior `nan_value_counts` counts only the elements of the rows written, matching iceberg-java. Compacting the list/map child arrays (or the whole batch) before the metrics are computed would fix it. ### Additional context Found in an audit of the native Iceberg write path. Part of #5649. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
