voonhous opened a new pull request, #20136: URL: https://github.com/apache/hudi/pull/20136
### Describe the issue this Pull Request addresses Fixes #20135. On Spark 4.1 with `spark.sql.variant.pushVariantIntoScan` on, `ALTER TABLE ... ADD COLUMNS (v2 variant)` without schema-on-read breaks every Hudi row-based parquet read of the files written before the DDL: pushed-down reads of `v2`, the COW update of a pre-DDL file group through the SPARK-record merge handle, and MOR compaction of such a group all fail with Spark's `INVALID_VARIANT_SHREDDING_SCHEMA`. Cause: Hudi's readers force `spark.sql.optimizer.nestedSchemaPruning.enabled` off, so Spark's `ParquetReadSupport` skips `intersectParquetGroups` and a requested column the file lacks stays in the parquet read schema as a group synthesised from its catalyst type. `HoodieParquetReadSupport.trimParquetSchema` already drops such nested fields but kept a missing top-level one. For a plain column parquet-mr null-fills it; for a variant column requested as the PushVariantIntoScan projection struct (or the internal reader's full-variant struct) the synthesised group is a plain struct, Spark still routes it to `ParquetVariantConverter`, and `buildVariantSchema` rejects it. Vanilla Spark 4.1 shows the same error with pruning off and returns nulls with pruning on. Found by item 4 of the `PushVariantIntoScan` validation checklist on #18285 (https://github.com/apache/hudi/issues/18285#issuecomment-5869224297). ### Summary and Changelog - `HoodieParquetReadSupport.trimParquetSchema` drops a top-level field the file does not have for the row-based reader, the way it already drops missing nested fields. `ParquetRowConverter` builds converters per parquet field and leaves the catalyst columns it never sees null, which is what Spark's own reader gets from `intersectParquetGroups`. The vectorized reader keeps the field, as in Spark; every Hudi row read constructs this read support with `enableVectorizedReader = false`. - `TestHoodieParquetReadSupport` pins both arms over a projection-shaped missing field. - `TestVariantShreddingMixedLayouts` gets "A second variant column added by DDL reads through old and new files": a shredded base with `v` only; the DDL pinned as an `ALTER_SCHEMA` commit whose schema types `v2` as VARIANT, with no data file touched; reads over the old file alone; an insert under a zero small-file limit that opens a second file group with `v2` shredded; an update of one row per group that rewrites both groups on COW (the old group's new base version carries `v2`, null for untouched rows) and lands as one data block per group on MOR, followed by compaction; every state read on both pushdown arms with the plan pinned per arm and per column. Legs: COW with both record types, MOR with native parquet log blocks and with avro data blocks on table version 9. ### Impact Row-based parquet reads of a file lacking a requested top-level column now return null for it on every Spark version, instead of relying on parquet-mr to synthesise the column from a converted type. Existing add-column reads of plain columns keep the same result. No config change. ### Risk Level low ### Documentation Update none ### Contributor's checklist - [x] Read through [contributor's guide](https://hudi.apache.org/contribute/how-to-contribute) - [x] Enough context is provided in the sections above - [x] Adequate tests were added if applicable -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
