ErikBPF commented on PR #5903: URL: https://github.com/apache/datafusion-comet/pull/5903#issuecomment-5951378757
Updated against pinned main `63fd1c9e`. The three original patches are unchanged. Published ancestry is retained with a merge commit whose tree is identical to the tested refreshed revision; no force-push was used. The fix recursively widens nullability for intermediate `collect_list`/`collect_set` buffers, without changing final result types. The regression covers structs and arrays of structs with required nested fields, compares Spark results, and requires native aggregation with actual spilling. Verification on Apollo: - Current-main, test-only baseline: the regression failed with the intended nested-nullability schema mismatch. - Spark 4.1 and 3.5: the fixed regression passed on both profiles. Full aggregate suites passed 129 and 127 tests respectively; two ignored in each profile and two existing version cancellations on 3.5. - Full Spark SQL `sql_core-1`: 8,341 passed, zero failed; 64 canceled, 580 ignored, no aborted suites. An isolated FHS environment supplied Spark's hard-coded `/bin/bash`; the host was not modified. SQL includes disabled-Comet control sessions, so the full SQL count is not a claim that every query ran natively. The regression separately verifies native execution and positive spill. Updated head: `b6b680b7f0245c414331b4d4388376854460f2c8` (tested tree unchanged from `06479c310b72abe3eec1d57461dbbcbe9fb1f3ee`). -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
