voonhous opened a new pull request, #20136:
URL: https://github.com/apache/hudi/pull/20136

   ### Describe the issue this Pull Request addresses
   
   Fixes #20135. On Spark 4.1 with `spark.sql.variant.pushVariantIntoScan` on, 
`ALTER TABLE ... ADD COLUMNS (v2 variant)` without schema-on-read breaks every 
Hudi row-based parquet read of the files written before the DDL: pushed-down 
reads of `v2`, the COW update of a pre-DDL file group through the SPARK-record 
merge handle, and MOR compaction of such a group all fail with Spark's 
`INVALID_VARIANT_SHREDDING_SCHEMA`.
   
   Cause: Hudi's readers force 
`spark.sql.optimizer.nestedSchemaPruning.enabled` off, so Spark's 
`ParquetReadSupport` skips `intersectParquetGroups` and a requested column the 
file lacks stays in the parquet read schema as a group synthesised from its 
catalyst type. `HoodieParquetReadSupport.trimParquetSchema` already drops such 
nested fields but kept a missing top-level one. For a plain column parquet-mr 
null-fills it; for a variant column requested as the PushVariantIntoScan 
projection struct (or the internal reader's full-variant struct) the 
synthesised group is a plain struct, Spark still routes it to 
`ParquetVariantConverter`, and `buildVariantSchema` rejects it. Vanilla Spark 
4.1 shows the same error with pruning off and returns nulls with pruning on.
   
   Found by item 4 of the `PushVariantIntoScan` validation checklist on #18285 
(https://github.com/apache/hudi/issues/18285#issuecomment-5869224297).
   
   ### Summary and Changelog
   
   - `HoodieParquetReadSupport.trimParquetSchema` drops a top-level field the 
file does not have for the row-based reader, the way it already drops missing 
nested fields. `ParquetRowConverter` builds converters per parquet field and 
leaves the catalyst columns it never sees null, which is what Spark's own 
reader gets from `intersectParquetGroups`. The vectorized reader keeps the 
field, as in Spark; every Hudi row read constructs this read support with 
`enableVectorizedReader = false`.
   - `TestHoodieParquetReadSupport` pins both arms over a projection-shaped 
missing field.
   - `TestVariantShreddingMixedLayouts` gets "A second variant column added by 
DDL reads through old and new files": a shredded base with `v` only; the DDL 
pinned as an `ALTER_SCHEMA` commit whose schema types `v2` as VARIANT, with no 
data file touched; reads over the old file alone; an insert under a zero 
small-file limit that opens a second file group with `v2` shredded; an update 
of one row per group that rewrites both groups on COW (the old group's new base 
version carries `v2`, null for untouched rows) and lands as one data block per 
group on MOR, followed by compaction; every state read on both pushdown arms 
with the plan pinned per arm and per column. Legs: COW with both record types, 
MOR with native parquet log blocks and with avro data blocks on table version 9.
   
   ### Impact
   
   Row-based parquet reads of a file lacking a requested top-level column now 
return null for it on every Spark version, instead of relying on parquet-mr to 
synthesise the column from a converted type. Existing add-column reads of plain 
columns keep the same result. No config change.
   
   ### Risk Level
   
   low
   
   ### Documentation Update
   
   none
   
   ### Contributor's checklist
   
   - [x] Read through [contributor's 
guide](https://hudi.apache.org/contribute/how-to-contribute)
   - [x] Enough context is provided in the sections above
   - [x] Adequate tests were added if applicable
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to