voonhous opened a new issue, #20140:
URL: https://github.com/apache/hudi/issues/20140

   ## Bug Description
   
   **What happened:**
   
   `HoodieFileGroupReaderBasedFileFormat.buildReaderWithPartitionValues` writes 
the per-query vectorization decision into the SESSION conf:
   
   ```scala
   
spark.sessionState.conf.setConfString("spark.sql.parquet.enableVectorizedReader",
 supportVectorizedRead.toString)
   ```
   
   
(`hudi-spark-datasource/hudi-spark-common/src/main/scala/org/apache/spark/sql/execution/datasources/parquet/HoodieFileGroupReaderBasedFileFormat.scala`,
 introduced by #14161.)
   
   `supportBatch` turns vectorization off for any query whose schema carries a 
variant (Spark 4.1+, #18605), for a vector column, and for 
MOR/incremental/bootstrap reads. Each such query therefore sets 
`spark.sql.parquet.enableVectorizedReader=false` for the whole session, and 
every later parquet read in that session, Hudi or not, runs row-based until 
some Hudi query with batch support flips it back to true. A user who set the 
conf explicitly gets it overwritten either way.
   
   **To reproduce** (Spark SQL, Spark 4.1):
   
   ```sql
   create table t (id int, v variant, ts long) using hudi
     tblproperties (primaryKey = 'id', preCombineField = 'ts') location 
'/tmp/t';
   insert into t values (1, parse_json('{"a":1}'), 1000);
   set spark.sql.parquet.enableVectorizedReader;   -- true
   select id, v from t;
   set spark.sql.parquet.enableVectorizedReader;   -- false
   ```
   
   **Effect on tests:**
   
   Whether a Hudi read is vectorized now depends on which query ran before it 
in the same session. `TestVariantShreddingMixedLayouts."Schema-on-read reads of 
shredded variant files fail fast"` pins `count(*)` after a schema-on-read DDL 
as passing; it passes only because the variant reads before it turned the 
session conf off, and fails when run first (#20139).
   
   **Expected behavior:**
   
   The batch decision should reach the base-file reader through the reader's 
own arguments (it already does: `buildBaseFileReader(..., 
supportVectorizedRead)`), not through a session conf that outlives the query.
   
   ## Environment
   
   Spark 4.1.1, Scala 2.13, JDK 17, master `9e9f7336a49e`. Found while probing 
rename under schema-on-read for #18285 (checklist item 5).
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to