andygrove opened a new issue, #6549:
URL: https://github.com/apache/datafusion-comet/issues/6549

   ### Describe the bug
   
   Since Spark 4.0, `ArrayBasedMapBuilder` normalizes floating-point map keys 
before it checks for duplicates. `-0.0` becomes `0.0` and NaN is canonicalized, 
unless `spark.sql.legacy.disableMapKeyNormalization` is set. Spark 3.4 and 3.5 
don't normalize map keys.
   
   Comet's native map construction doesn't fully match this. With the Spark 4.1 
profile and default configs on `branch-1.1`:
   
   - `map_from_entries` keeps a `-0.0` key, where Spark returns `0.0`.
   - `map_from_entries` and `map_from_arrays` both accept `0.0` and `-0.0` as 
two keys in the same map, where Spark raises `DUPLICATED_MAP_KEY`.
   - `map_from_arrays` does normalize a lone `-0.0` key, so only its duplicate 
check differs.
   
   The wrong keys flow into anything downstream, for example 
`explode(map_keys(...))`. `CometMapFromArrays` and `CometMapFromEntries` report 
`Compatible` for float and double keys, so this runs natively by default. Map 
lookups on float keys already fall back for this reason (#5580), but map 
construction doesn't.
   
   ### Steps to reproduce
   
   ```sql
   -- Spark 4.0, 4.1 or 4.2, default configs
   CREATE TABLE fmk_neg(k ARRAY<DOUBLE>, v ARRAY<INT>) USING parquet;
   INSERT INTO fmk_neg VALUES (array(CAST('-0.0' AS DOUBLE), 1.5D), array(10, 
20));
   
   SELECT map_from_entries(arrays_zip(k, v)) FROM fmk_neg;
   -- Spark: {0.0 -> 10, 1.5 -> 20}
   -- Comet: {-0.0 -> 10, 1.5 -> 20}
   
   SELECT key FROM fmk_neg LATERAL VIEW 
explode(map_keys(map_from_entries(arrays_zip(k, v)))) e AS key;
   -- Spark: 0.0, 1.5
   -- Comet: -0.0, 1.5
   
   CREATE TABLE fmk_dup(k ARRAY<DOUBLE>, v ARRAY<INT>) USING parquet;
   INSERT INTO fmk_dup VALUES (array(0.0D, CAST('-0.0' AS DOUBLE)), array(1, 
2));
   
   SELECT map_from_entries(arrays_zip(k, v)) FROM fmk_dup;
   SELECT map_from_arrays(k, v) FROM fmk_dup;
   -- Spark: DUPLICATED_MAP_KEY for both
   -- Comet: both return a map
   ```
   
   I ran these as SQL file tests on `branch-1.1` at 1de7e7d with the default 
Spark 4.1.3 profile. `SELECT map_from_arrays(k, v) FROM fmk_neg`, and 
`map_concat` over it, matched Spark.
   
   ### Expected behavior
   
   On Spark 4.0+, match Spark: normalize float and double keys before the 
duplicate check, honoring `spark.sql.legacy.disableMapKeyNormalization`. Until 
the native kernels do that, fall back to Spark for float and double keys on 
Spark 4.0+.
   
   ### Additional context
   
   - I found this while reviewing #6107, which makes map `explode` native and 
so adds one more consumer of these keys. The bug predates that PR.
   - `create_map` with float keys hasn't been checked.
   - Related to #6385.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to