sunchao opened a new pull request, #6428:
URL: https://github.com/apache/datafusion-comet/pull/6428
## Which issue does this PR close?
No associated issue.
## Rationale for this change
Spark compares struct fields by position, but native nested-type
reconciliation merged them by name. On the PR base, this query succeeds in
Spark and fails in Comet with `Nested predicate requires matching types`:
```sql
SELECT named_struct('x', CAST(id AS DOUBLE), 'y', CAST(NULL AS DOUBLE)) =
named_struct('y', CAST(NULL AS DOUBLE), 'x', CAST(id AS DOUBLE))
FROM range(8)
```
Native `named_struct` also inferred field nullability from native children,
which can be more conservative than Catalyst. Keeping Catalyst's declared flags
and reconciling compatible nested types by position gives arrays, comparisons,
and conditional branches consistent Arrow types.
## What changes are included in this PR?
- Serialize Catalyst field nullability for `CreateNamedStruct`, validate
metadata arity, and retain the flags through type inference, evaluation, and
child rewrites.
- Reconcile compatible nested IF/CASE branches and comparison operands by
position, using Comet's Spark-compatible cast while retaining general CASE type
promotion for other layouts.
- Add native coverage for typed NULL structs, dictionary children, metadata
arity, and uniform/mixed IF selection, plus SQL regressions that require native
execution for arrays, CASE, IF, and comparisons.
## How are these changes tested?
- On base `aca67fd581fb1838a2745f5676b49f4deee3bb8d`, the new SQL fixture
fails on the comparison above. The earlier array and CASE queries pass on that
base.
- Four focused native constructor tests pass, including the
nullability/array regression, metadata arity, and existing dictionary-child
controls.
- The native IF regression passes its 36 combinations of nested types,
branch order, and uniform/mixed predicates.
- Root-reactor Spark 4.1 Spotless and changed-source Rust formatting pass,
as does `git diff --check`.
- On the patched head, the new SQL fixture passes all six queries with Spark
4.1.3, including the comparison that fails on the base. All 14 existing
conditional SQL test variants also pass.
- The broader Spark SQL suite will be requested with `run-spark-4.1-tests`.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]