andygrove opened a new issue, #6546:
URL: https://github.com/apache/datafusion-comet/issues/6546
### Describe the bug
When a field inside a struct column, or inside the struct of an array, is
renamed with `ALTER TABLE ... RENAME COLUMN`, the native Iceberg scan returns
NULL for that field in rows from data files written before the rename. There is
no error, and Spark returns the values.
This is a regression in 1.1.0. 1.0.0 returns the right values for the
queries below.
iceberg-rust compares the file's nested type with the table's using arrow's
`equals_datatype`, which ignores field names. When the two match, it passes the
column through. The batch is labeled with the table's field names, but the
struct array keeps the file's old names (apache/iceberg-rust#2617). In 1.0.0,
`CometCastColumnExpr` relabeled such a column by position to the names Spark
asked for. In 1.1.0 the column goes through DataFusion's struct cast instead,
which matches fields by the array's names and fills the renamed field with
NULL. This most likely dates from #5262, which added
`is_pure_structural_narrowing` to keep DataFusion's `CastExpr` when the
declared struct field names match.
The bug shows up when the file's nested fields have the same nullability as
the table's, which is the usual case. Iceberg writes a file with the
nullability of the inserted data, so any null inside the struct makes the
fields optional, as the table declares them. Spark 4 writes a struct built only
from non-null literals with required fields. iceberg-rust then casts the column
by position instead, and a plain rename reads correctly. That is why the
reproduction below includes a row with nulls. Spark 3.4 writes the fields as
optional either way.
### Steps to reproduce
```sql
CREATE TABLE cat.db.evo_rename (id INT, s STRUCT<a: INT, b: STRING>, items
ARRAY<STRUCT<a: INT, b: INT>>) USING iceberg;
INSERT INTO cat.db.evo_rename VALUES
(1, named_struct('a', 1, 'b', 'x'), array(named_struct('a', 1, 'b', 2))),
(2, named_struct('a', CAST(NULL AS INT), 'b', CAST(NULL AS STRING)),
array(CAST(NULL AS STRUCT<a: INT, b: INT>)));
ALTER TABLE cat.db.evo_rename RENAME COLUMN s.a TO z;
ALTER TABLE cat.db.evo_rename RENAME COLUMN items.element.a TO z;
SELECT id, s FROM cat.db.evo_rename ORDER BY id;
SELECT id, items FROM cat.db.evo_rename ORDER BY id;
```
With the native Iceberg scan on, which is the default, 1.1.0-rc1 returns
`(1, {null, x})` for the first query and `(1, [{null, 2}])` for the second.
### Expected behavior
Spark's results, which 1.0.0 also returns: `(1, {1, x})` and `(2, {null,
null})` for the first query, and `(1, [{1, 2}])` and `(2, [null])` for the
second.
### Workaround
`spark.comet.scan.icebergNative.enabled=false` reads every Iceberg table
with Spark's reader.
### Additional context
Two related cases were already wrong in 1.0.0:
- Reading only the renamed field, `SELECT id, s.z`, returned NULL in 1.0.0.
In 1.1.0 it fails with `Cannot cast struct with 2 fields to 1 fields because
there is no field name overlap`.
- A reorder plus a rename, `ALTER COLUMN m.b FIRST` followed by `RENAME
COLUMN m.a TO c` on `m STRUCT<a: INT, b: INT>`, returns `{1, 2}` in 1.0.0 and
`{2, null}` in 1.1.0, where Spark returns `{2, 1}`.
Verified on Spark 4.1 with Iceberg 1.11.0, comparing the 1.0.0 tag
(3a7a2c437c) with 1.1.0-rc1 (470fc7826b). The same NULLs appear on `main` on
Spark 3.4 and 4.1.
Found while fixing #6504. #6543 makes the native scan fall back when a
projected column has a nested field that was added or renamed over the table's
schema history, so it fixes this too. apache/iceberg-rust#3255 reconciles
nested fields by field id, which would fix the iceberg-rust side.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]