voonhous commented on issue #20135: URL: https://github.com/apache/hudi/issues/20135#issuecomment-5889758469
Some notes for myself when referring back to this fix in the future: - Spark's read support, pruning on. `ParquetReadSupport` intersects the projection with the file schema, so a missing column never reaches `parquet-mr`. The row converter starts every record with all fields null and only fills the fields it has a parquet column for. That is the null-fill you are thinking of. Hudi switches nested schema pruning off, so this step is skipped in Hudi's readers. - `parquet-mr` itself, pruning off. The missing column stays in the projection, with a parquet type synthesised from the catalyst type. `parquet-mr` accepts a projection with an optional column the file lacks and returns nulls for it. This is why adding a plain string column has always worked in Hudi. The second path has a precondition the first does not: Spark must be able to build a converter for the synthesised column before a single row is read. For a string that is trivial. For a variant requested as the pushdown projection struct, the synthesised type is a plain group, Spark picks its variant converter because the catalyst type is variant-shaped, and that converter's constructor validates the parquet group against the variant shredding spec and throws. The failure happens while the reader is being prepared, so the null-fill mechanism is never reached. The fix is to trim away missing variant column in the requested schema if it sees one when reading an older parquet with an older schema (prior to the schema evolution of adding a Variant v2 column). -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
