andygrove opened a new issue, #42: URL: https://github.com/apache/datafusion-iceberg/issues/42
### Describe the bug `IcebergTableProvider` and `IcebergStaticTableProvider::try_new_from_table` declare the table's current schema ([table/mod.rs#L94-L97](https://github.com/apache/datafusion-iceberg/blob/a2bc9427d0659591b5f8122b90fec710fe2f5de6/crates/datafusion/src/table/mod.rs#L94-L97), [table/mod.rs#L258-L262](https://github.com/apache/datafusion-iceberg/blob/a2bc9427d0659591b5f8122b90fec710fe2f5de6/crates/datafusion/src/table/mod.rs#L258-L262)). iceberg-rust's scan binds the projection and filters to the schema of the snapshot it reads instead: `TableScanBuilder::build` uses `snapshot.schema(...)`. After a schema change that has no snapshot of its own, the provider declares columns the scan can't find. ### To Reproduce 1. Create `t(a INT NOT NULL)` and insert one row. 2. Add an optional column `c STRING`, for example with `Transaction::update_schema().add_column(AddColumn::optional("c", ...))` or with `ALTER TABLE ... ADD COLUMN` from Spark. Don't write any new data. 3. Build a fresh provider and run `SELECT * FROM t`: ``` DataInvalid => Column c not found in table. Schema: table { 1: a: required int } ``` ### Expected behavior The query returns `a = 1, c = NULL`. ### Additional context The root cause is the iceberg-rust scan behavior tracked in apache/iceberg-rust#2565 and apache/iceberg-rust#2905. Until that's fixed upstream, this crate could declare the scanned snapshot's schema, or fill columns missing from it with nulls. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
