voonhous opened a new pull request, #20125: URL: https://github.com/apache/hudi/pull/20125
### Describe the issue this Pull Request addresses Validation of Spark 4.1+ `PushVariantIntoScan` on master, schema-evolution item 1 of the checklist on #18285 (https://github.com/apache/hudi/issues/18285#issuecomment-5869224297): implicit widening of a sibling column beside a nested variant on MOR tables with log files. #19783 fixed `SparkSchemaTransformUtils.addMissingFields` folding the PushVariantIntoScan projection struct at `s.inner` back to `VariantType` when a sibling member of `s` had an implicit type change (file `s.n` int, table long), and pinned it on COW only. On MOR the file-group reader merges rows from the base file and from log blocks, and each source reconciles its own file schema against the widened requested schema before the merge, so the row shapes have to agree. That arm had no test. ### Summary and Changelog Test only. `TestVariantShreddingMixedLayouts` gets the MOR twin of the COW leg, "Implicit widening of a sibling keeps the nested variant projection through MOR log blocks": - int base file (ids 0-2), an int log block written before the widening by a SQL update (ids 1-2), then one DataFrame upsert with `s.n` as bigint that widens the table schema, appends a bigint log block for id 2 and opens a second file group with a bigint base file for ids 3-4; - every slot read back through `variant_get`, one filter per slot, the widened sibling itself and a whole-value cast, with `spark.sql.variant.pushVariantIntoScan` on and off and the plan pinned per arm; - two legs: native parquet log blocks (SPARK record type, current table version) and avro data blocks (AVRO record type on table version 9, with the DataFrame write pinning `hoodie.write.table.version` so auto-upgrade does not switch the table to native logs); - the physical layout is asserted: two base files, and the data blocks' `s.n` types are exactly int and long. Result: green on master as-is. No production change. ### Impact None. ### Risk Level none ### Documentation Update none ### Contributor's checklist - [x] Read through [contributor's guide](https://hudi.apache.org/contribute/how-to-contribute) - [x] Enough context is provided in the sections above - [x] Adequate tests were added if applicable -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
