comphead opened a new pull request, #6326:
URL: https://github.com/apache/datafusion-comet/pull/6326
## Which issue does this PR close?
Closes #6136.
## Rationale for this change
Under `spark.sql.caseSensitive=false`, Spark's Parquet reader
(`ParquetReadSupport.clipParquetGroupFields`) matches each requested field to
the file fields sharing its `toLowerCase(Locale.ROOT)` name, and raises `Found
duplicate field(s) "x": [x, X] in case-insensitive mode` when more than one
answers. The native scan returned rows instead.
Comet resolves nested names only while casting a column whose file type
differs from the requested type. `check_conversion` and `CometCastColumnExpr`
both go through `match_struct_fields`. When the two types are equal, nothing
resolves the names:
- DataFusion's `DefaultPhysicalExprAdapter` leaves a column as a bare
`Column` whenever its logical and physical `Field`s compare equal.
- The Parquet opener skips the adapter altogether when the file schema
equals the logical schema and no predicate is pushed. That is the metadata-free
case the issue describes.
So `s struct<x, X>` read from a file with the same shape came back
positionally.
One correction to the issue: Spark's analyzer does check nested names. On
3.4 through 4.1, `DataSource.resolveRelation` runs
`SchemaUtils.checkSchemaColumnNameDuplication` recursively under the session
resolver, so `spark.read.schema("s struct<x: bigint, X: bigint>")` fails with
`COLUMN_ALREADY_EXISTS` under the case-insensitive resolver. The shape still
reaches the scan when a DataFrame is analyzed under the case-sensitive resolver
and planned under the case-insensitive one. The same flip on top-level columns
never reaches Comet, because `FileSourceStrategy` fails with
`AMBIGUOUS_REFERENCE` while resolving the relation's output.
The colliding names are in the requested schema, so this is decidable from
the plan, as #6004 did for repeated field ids.
## What changes are included in this PR?
1. `DataTypeSupport.hasCaseInsensitiveDuplicateFieldNames` is true when two
sibling fields anywhere in a type fold to one name under
`toLowerCase(Locale.ROOT)`, the fold Spark's reader uses. It walks struct
children at every depth, including array elements and map keys and values.
2. `CometNativeScan.isSupported` declines the scan when
`spark.sql.caseSensitive=false` and the required schema trips that predicate,
so Spark's reader resolves it and reports the ambiguity. Only the V1 native
scan is gated. The Iceberg scan resolves by field id and is unchanged.
3. The Parquet compatibility guide lists the new fallback.
Like #6004, the check reads the requested schema, not the file. It can
decline a read Spark would accept, such as a file without the colliding
sibling, or siblings that resolve by field id. That costs native execution but
not correctness, and such schemas only reach the scan past the analyzer check
above. No Parquet decoding changes.
The rest of this family was already covered. Byte-identical duplicate names
(#5866) and repeated field ids (#6004) in the requested schema are declined at
planning time. A duplicate present only in the file makes the file type differ
from the requested type, so the native checks run.
## How are these changes tested?
- `CometNativeReaderSuite`: end to end, with the issue's reproduction. A
file written without key-value metadata, with schema `s struct<x, X>`, is read
back with that same schema.
- Analyzed and planned under the case-sensitive resolver, the scan stays
native and matches Spark.
- Analyzed under the case-sensitive resolver, then planned and run under
the case-insensitive one, the plan has no `CometNativeScanExec`, carries the
new fallback reason, and raises Spark's `Found duplicate field(s) "x": [x, X]
in case-insensitive mode`.
- `CometScanRuleSuite`: the predicate detects collisions at the root, in
nested and deeper structs, in array elements, and in map keys and values,
including non-ASCII (`CAFÉ` / `café`) and byte-identical names. It does not
flag the same name in separate groups, or a parent and child that share a name.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]