comphead opened a new pull request, #6326:
URL: https://github.com/apache/datafusion-comet/pull/6326

   ## Which issue does this PR close?
   
   Closes #6136.
   
   ## Rationale for this change
   
   Under `spark.sql.caseSensitive=false`, Spark's Parquet reader 
(`ParquetReadSupport.clipParquetGroupFields`) matches each requested field to 
the file fields sharing its `toLowerCase(Locale.ROOT)` name, and raises `Found 
duplicate field(s) "x": [x, X] in case-insensitive mode` when more than one 
answers. The native scan returned rows instead.
   
   Comet resolves nested names only while casting a column whose file type 
differs from the requested type. `check_conversion` and `CometCastColumnExpr` 
both go through `match_struct_fields`. When the two types are equal, nothing 
resolves the names:
   
   - DataFusion's `DefaultPhysicalExprAdapter` leaves a column as a bare 
`Column` whenever its logical and physical `Field`s compare equal.
   - The Parquet opener skips the adapter altogether when the file schema 
equals the logical schema and no predicate is pushed. That is the metadata-free 
case the issue describes.
   
   So `s struct<x, X>` read from a file with the same shape came back 
positionally.
   
   One correction to the issue: Spark's analyzer does check nested names. On 
3.4 through 4.1, `DataSource.resolveRelation` runs 
`SchemaUtils.checkSchemaColumnNameDuplication` recursively under the session 
resolver, so `spark.read.schema("s struct<x: bigint, X: bigint>")` fails with 
`COLUMN_ALREADY_EXISTS` under the case-insensitive resolver. The shape still 
reaches the scan when a DataFrame is analyzed under the case-sensitive resolver 
and planned under the case-insensitive one. The same flip on top-level columns 
never reaches Comet, because `FileSourceStrategy` fails with 
`AMBIGUOUS_REFERENCE` while resolving the relation's output.
   
   The colliding names are in the requested schema, so this is decidable from 
the plan, as #6004 did for repeated field ids.
   
   ## What changes are included in this PR?
   
   1. `DataTypeSupport.hasCaseInsensitiveDuplicateFieldNames` is true when two 
sibling fields anywhere in a type fold to one name under 
`toLowerCase(Locale.ROOT)`, the fold Spark's reader uses. It walks struct 
children at every depth, including array elements and map keys and values.
   2. `CometNativeScan.isSupported` declines the scan when 
`spark.sql.caseSensitive=false` and the required schema trips that predicate, 
so Spark's reader resolves it and reports the ambiguity. Only the V1 native 
scan is gated. The Iceberg scan resolves by field id and is unchanged.
   3. The Parquet compatibility guide lists the new fallback.
   
   Like #6004, the check reads the requested schema, not the file. It can 
decline a read Spark would accept, such as a file without the colliding 
sibling, or siblings that resolve by field id. That costs native execution but 
not correctness, and such schemas only reach the scan past the analyzer check 
above. No Parquet decoding changes.
   
   The rest of this family was already covered. Byte-identical duplicate names 
(#5866) and repeated field ids (#6004) in the requested schema are declined at 
planning time. A duplicate present only in the file makes the file type differ 
from the requested type, so the native checks run.
   
   ## How are these changes tested?
   
   - `CometNativeReaderSuite`: end to end, with the issue's reproduction. A 
file written without key-value metadata, with schema `s struct<x, X>`, is read 
back with that same schema.
     - Analyzed and planned under the case-sensitive resolver, the scan stays 
native and matches Spark.
     - Analyzed under the case-sensitive resolver, then planned and run under 
the case-insensitive one, the plan has no `CometNativeScanExec`, carries the 
new fallback reason, and raises Spark's `Found duplicate field(s) "x": [x, X] 
in case-insensitive mode`.
   - `CometScanRuleSuite`: the predicate detects collisions at the root, in 
nested and deeper structs, in array elements, and in map keys and values, 
including non-ASCII (`CAFÉ` / `café`) and byte-identical names. It does not 
flag the same name in separate groups, or a parent and child that share a name.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to