dwsmith1983 commented on PR #25895: URL: https://github.com/apache/datafusion/pull/25895#issuecomment-5922233473
> Would you mind adding more information how critical is not having this fix in the patchset? Without it the query returns wrong results and raises no error. With `coerce_int96` on, struct, list and map fields read from a file that has an INT96 column lose their metadata. A reader that matches columns by Parquet field id then no longer finds those containers and fills them with nulls. Comet enables the coercion for every scan, so with Spark's field id reads on, `where s.inner.x = 1` and `where s is not null` return no rows and `count(s)` is 0 where Spark reads the data (apache/datafusion-comet#6405). Until DataFusion carries this, Comet's stopgap (apache/datafusion-comet#6444) sends every scan that asks for an id on a nested field back to Spark. The fix is three `.with_metadata(...)` calls in `schema_coercion.rs`. Files without an INT96 column are unaffected. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
