voonhous opened a new issue, #20139:
URL: https://github.com/apache/hudi/issues/20139
## Bug Description
**What happened:**
On a table with a VARIANT column read with
`hoodie.schema.on.read.enable=true`, `select count(*)` fails once any
schema-on-read DDL has committed an internal schema (an `add columns` is
enough). Spark 4.1.1, master `9e9f7336a49e`, COW, SPARK record type, shredded
or unshredded file:
```
org.apache.spark.SparkException: [FAILED_READ_FILE.NO_HINT] Encountered
error while reading file ...
Caused by: java.lang.NullPointerException
at
org.apache.spark.sql.execution.datasources.parquet.HoodieVectorizedParquetRecordReader.close(HoodieVectorizedParquetRecordReader.java:97)
at
org.apache.spark.sql.execution.datasources.RecordReaderIterator.close(RecordReaderIterator.scala:69)
at
org.apache.spark.sql.execution.datasources.parquet.Spark41ParquetReader.buildVectorizedIterator(Spark41ParquetReader.scala:251)
```
The NPE masks the real failure. With a null guard on `close`, the exception
is Spark's:
```
org.apache.spark.sql.AnalysisException: Invalid Spark read type: expected
optional group v (VARIANT(1)) { required binary metadata; optional binary
value; } to be variant type but found STRUCT<metadata: BINARY NOT NULL, value:
BINARY NOT NULL>
at
org.apache.spark.sql.execution.datasources.parquet.ParquetSchemaConverter$.checkConversionRequirement(ParquetSchemaConverter.scala:897)
at
org.apache.spark.sql.execution.datasources.parquet.ParquetToSparkSchemaConverter.convertGroupField(ParquetSchemaConverter.scala:383)
...
at
org.apache.spark.sql.execution.datasources.parquet.VectorizedParquetRecordReader.initialize(VectorizedParquetRecordReader.java:199)
```
**To reproduce** (Spark SQL, Spark 4.1):
```sql
create table t (id int, v variant, ts long) using hudi
tblproperties (primaryKey = 'id', preCombineField = 'ts', type = 'cow')
location '/tmp/t';
insert into t values (1, parse_json('{"a":1}'), 1000);
set hoodie.schema.on.read.enable=true;
alter table t add columns (note string);
select count(*) from t; -- fails as above in a fresh session
```
**Cause:**
Two parts.
1. `ParquetSchemaEvolutionUtils.getHadoopConfClone` skips the
shredded-variant guard for an empty projection (`requiredSchema.isEmpty`) but
still merges the file schema with the UNPRUNED query schema and sets the result
as `SPARK_ROW_REQUESTED_SCHEMA`. The internal schema has no VARIANT arm, so the
variant column arrives as `struct<metadata, value>`. Spark 4.1's vectorized
reader (`ParquetToSparkSchemaConverter.checkConversionRequirement`) refuses a
VARIANT-annotated parquet group whose requested Spark type is not variant. The
row-based reader accepts the same request
(`spark.sql.parquet.enableVectorizedReader=false` makes the count pass).
2. `HoodieVectorizedParquetRecordReader.close` iterates `idToColumnVectors`,
which `initBatch` creates. `Spark41ParquetReader.buildVectorizedIterator`
closes the iterator when `initialize` throws, before `initBatch` ran, so the
close NPEs and the original exception is lost.
**Why the existing test does not catch it:**
`TestVariantShreddingMixedLayouts."Schema-on-read reads of shredded variant
files fail fast"` pins `count(*)` after the DDL and is green, but only because
the two variant reads before it made
`HoodieFileGroupReaderBasedFileFormat.buildReaderWithPartitionValues` write
`spark.sql.parquet.enableVectorizedReader=false` into the session conf (see the
companion issue on that session-conf write). Run the count as the first read on
a fresh session and it fails.
**Expected behavior:**
An empty projection reads no column data, so the requested parquet schema
for it should be empty (or pruned to the columns actually read), not the whole
merged table schema; and `close` should tolerate a reader whose `initBatch`
never ran so the real exception surfaces.
## Environment
Spark 4.1.1, Scala 2.13, JDK 17, local filesystem, master `9e9f7336a49e`.
Found while probing rename under schema-on-read for #18285 (checklist item 5).
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]