voonhous opened a new pull request, #20039:
URL: https://github.com/apache/hudi/pull/20039

   ### Describe the issue this Pull Request addresses
   
   Closes #20037. Supersedes the do-not-merge evidence PR #20034, whose suite 
lands here as the regression test.
   
   release-1.2.1 carries #18674 but not #19783, so under Spark 4.1's default 
`spark.sql.variant.pushVariantIntoScan=true` a MOR read that merges a log file 
loses the row of a variant nulled through the log (`where v is null`), and 
fails the task on a struct-nested variant (`SparkOutOfMemoryError` on x86, 
SIGBUS on aarch64): the reader context overlaid and projected the pushed-down 
struct at the top level only, and the projector had no null guard.
   
   ### Summary and Changelog
   
   Hand-port of #19783 (master `c59987a024cd`); a cherry-pick conflicts in 10 
of 13 files because this branch has no mixed-layouts suite, no Spark4_2Adapter, 
and its reader context sits on the #18674 buffer-hook design. The CDC and 
MergeOnReadRDDV2 hunks of #19783 are comment-only and are not ported.
   
   - `SparkFileFormatInternalRowReaderContext`: the projection overlay and its 
detection recurse into struct members, mirroring PushVariantIntoScan's 
`VariantInRelation.rewriteType`; arrays and maps stay native.
   - `SparkAdapter`: `containsVariantProjection`, the shared walk the overlay 
and the projector key off.
   - `Spark4_1Adapter.buildVariantProjector`: recursive, and a null variant 
projects to a NULL struct instead of a struct of nulls.
   - `SparkSchemaTransformUtils`: the projection-over-VariantType arm, applied 
verbatim.
   - Tests: the two `TestBaseSpark4AdapterVariantMethods` cases from #19783, 
applied verbatim, plus `TestVariantPushVariantIntoScan` from #20034 (16 legs: 
COW/MOR x conf on/off x AVRO/SPARK record types, top-level and nested), now 
expected green.
   
   ### Impact
   
   Spark 4.1 MOR reads of variant columns on this branch return the right rows 
with the default conf. No API or format change.
   
   ### Risk Level
   
   low. Read path only, same code as master since 2026-09-01. Verified locally 
on Spark 4.1.1, JDK 17: `TestVariantPushVariantIntoScan` 16 of 16 legs pass 
(the 4 that failed on #20034's CI run are green), 
`TestBaseSpark4AdapterVariantMethods` 18 of 18, `TestVariantDataType` 9 of 9. 
Scalastyle: no new findings (hudi-spark-client carries pre-existing ones and is 
not style-checked in CI; hudi-spark4.1.x and hudi-spark are clean).
   
   ### Documentation Update
   
   none
   
   ### Contributor's checklist
   
   - [x] Read through [contributor's 
guide](https://hudi.apache.org/contribute/how-to-contribute)
   - [x] Enough context is provided in the sections above
   - [x] Adequate tests were added if applicable
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to