andygrove opened a new issue, #6546:
URL: https://github.com/apache/datafusion-comet/issues/6546

   ### Describe the bug
   
   When a field inside a struct column, or inside the struct of an array, is 
renamed with `ALTER TABLE ... RENAME COLUMN`, the native Iceberg scan returns 
NULL for that field in rows from data files written before the rename. There is 
no error, and Spark returns the values.
   
   This is a regression in 1.1.0. 1.0.0 returns the right values for the 
queries below.
   
   iceberg-rust compares the file's nested type with the table's using arrow's 
`equals_datatype`, which ignores field names. When the two match, it passes the 
column through. The batch is labeled with the table's field names, but the 
struct array keeps the file's old names (apache/iceberg-rust#2617). In 1.0.0, 
`CometCastColumnExpr` relabeled such a column by position to the names Spark 
asked for. In 1.1.0 the column goes through DataFusion's struct cast instead, 
which matches fields by the array's names and fills the renamed field with 
NULL. This most likely dates from #5262, which added 
`is_pure_structural_narrowing` to keep DataFusion's `CastExpr` when the 
declared struct field names match.
   
   The bug shows up when the file's nested fields have the same nullability as 
the table's, which is the usual case. Iceberg writes a file with the 
nullability of the inserted data, so any null inside the struct makes the 
fields optional, as the table declares them. Spark 4 writes a struct built only 
from non-null literals with required fields. iceberg-rust then casts the column 
by position instead, and a plain rename reads correctly. That is why the 
reproduction below includes a row with nulls. Spark 3.4 writes the fields as 
optional either way.
   
   ### Steps to reproduce
   
   ```sql
   CREATE TABLE cat.db.evo_rename (id INT, s STRUCT<a: INT, b: STRING>, items 
ARRAY<STRUCT<a: INT, b: INT>>) USING iceberg;
   INSERT INTO cat.db.evo_rename VALUES
     (1, named_struct('a', 1, 'b', 'x'), array(named_struct('a', 1, 'b', 2))),
     (2, named_struct('a', CAST(NULL AS INT), 'b', CAST(NULL AS STRING)), 
array(CAST(NULL AS STRUCT<a: INT, b: INT>)));
   ALTER TABLE cat.db.evo_rename RENAME COLUMN s.a TO z;
   ALTER TABLE cat.db.evo_rename RENAME COLUMN items.element.a TO z;
   SELECT id, s FROM cat.db.evo_rename ORDER BY id;
   SELECT id, items FROM cat.db.evo_rename ORDER BY id;
   ```
   
   With the native Iceberg scan on, which is the default, 1.1.0-rc1 returns 
`(1, {null, x})` for the first query and `(1, [{null, 2}])` for the second.
   
   ### Expected behavior
   
   Spark's results, which 1.0.0 also returns: `(1, {1, x})` and `(2, {null, 
null})` for the first query, and `(1, [{1, 2}])` and `(2, [null])` for the 
second.
   
   ### Workaround
   
   `spark.comet.scan.icebergNative.enabled=false` reads every Iceberg table 
with Spark's reader.
   
   ### Additional context
   
   Two related cases were already wrong in 1.0.0:
   
   - Reading only the renamed field, `SELECT id, s.z`, returned NULL in 1.0.0. 
In 1.1.0 it fails with `Cannot cast struct with 2 fields to 1 fields because 
there is no field name overlap`.
   - A reorder plus a rename, `ALTER COLUMN m.b FIRST` followed by `RENAME 
COLUMN m.a TO c` on `m STRUCT<a: INT, b: INT>`, returns `{1, 2}` in 1.0.0 and 
`{2, null}` in 1.1.0, where Spark returns `{2, 1}`.
   
   Verified on Spark 4.1 with Iceberg 1.11.0, comparing the 1.0.0 tag 
(3a7a2c437c) with 1.1.0-rc1 (470fc7826b). The same NULLs appear on `main` on 
Spark 3.4 and 4.1.
   
   Found while fixing #6504. #6543 makes the native scan fall back when a 
projected column has a nested field that was added or renamed over the table's 
schema history, so it fixes this too. apache/iceberg-rust#3255 reconciles 
nested fields by field id, which would fix the iceberg-rust side.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to