wombatu-kun commented on code in PR #20136:
URL: https://github.com/apache/hudi/pull/20136#discussion_r4140463292


##########
hudi-client/hudi-spark-client/src/main/scala/org/apache/spark/sql/execution/datasources/parquet/HoodieParquetReadSupport.scala:
##########
@@ -49,7 +49,11 @@ class HoodieParquetReadSupport(
     } else {
       readContext.getRequestedSchema
     }
-    val trimmedParquetSchema = 
HoodieParquetReadSupport.trimParquetSchema(requestedParquetSchema, 
context.getFileSchema)
+    // Same condition as Spark's own intersectParquetGroups: only the 
row-based reader wants a
+    // requested schema restricted to what the file has; the vectorized reader 
null-fills a missing
+    // column itself and matches columns by position.
+    val trimmedParquetSchema = 
HoodieParquetReadSupport.trimParquetSchema(requestedParquetSchema,
+      context.getFileSchema, dropMissingTopLevelFields = 
!enableVectorizedReader)

Review Comment:
   `HoodieSparkParquetReader.getUnsafeRowIterator` builds this read support 
with `enableVectorizedReader = true` for its row-based `ParquetReader`, so a 
parquet log block written before the `ADD COLUMNS` still keeps the missing `v2` 
and a pushed-down read of it should hit the same 
`INVALID_VARIANT_SHREDDING_SCHEMA`. No vectorized read ever goes through 
`HoodieParquetReadSupport` (every reader pins `READ_SUPPORT_CLASS` to Spark's 
`ParquetReadSupport`), so could the top-level drop be unconditional like the 
nested one, with an update before the DDL in the parquet-log MOR leg to cover 
it?



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to