voonhous commented on issue #20135:
URL: https://github.com/apache/hudi/issues/20135#issuecomment-5889758469

   Some notes for myself when referring back to this fix in the future:
   
   - Spark's read support, pruning on. `ParquetReadSupport` intersects the 
projection with the file schema, so a missing column never reaches 
`parquet-mr`. The row converter starts every record with all fields null and 
only fills the fields it has a parquet column for. That is the null-fill you 
are thinking of. Hudi switches nested schema pruning off, so this step is 
skipped in Hudi's readers.
   - `parquet-mr` itself, pruning off. The missing column stays in the 
projection, with a parquet type synthesised from the catalyst type. 
`parquet-mr` accepts a projection with an optional column the file lacks and 
returns nulls for it. This is why adding a plain string column has always 
worked in Hudi.
   
   The second path has a precondition the first does not: 
   Spark must be able to build a converter for the synthesised column before a 
single row is read. 
   
   For a string that is trivial. 
   
   For a variant requested as the pushdown projection struct, the synthesised 
type is a plain group, Spark picks its variant converter because the catalyst 
type is variant-shaped, and that converter's constructor validates the parquet 
group against the variant shredding spec and throws. The failure happens while 
the reader is being prepared, so the null-fill mechanism is never reached.
   
   The fix is to trim away missing variant column in the requested schema if it 
sees one when reading an older parquet with an older schema (prior to the 
schema evolution of adding a Variant v2 column).


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to