voonhous opened a new issue, #20141: URL: https://github.com/apache/hudi/issues/20141
## Summary Renaming a VARIANT column does not work today on either schema path. This issue lists every failure on the way, with its error and what has to change, so that the pins landing under #18285 (checklist item 5, https://github.com/apache/hudi/issues/18285#issuecomment-5869224297) reference one place instead of explaining the mechanism in test comments. Everything below was run on Spark 4.1.1, Scala 2.13, JDK 17, master `9e9f7336a49e`, COW and MOR, SPARK record type, shredded and unshredded base files, table `(id int, v variant, ts long)` renamed `v` to `w`. ## Schema-on-read (the only path with a rename primitive) **1. The DDL is refused under default configs.** ``` org.apache.hudi.exception.MissingSchemaFieldException: Schema validation failed due to missing field. Fields missing from incoming schema: {v} ``` `AlterTableCommand.commitWithSchema` runs `HoodieTable.validateSchema` since #13595, and the writer-schema check sees the rename as a dropped field. Tracked by #19766. `hoodie.datasource.write.schema.allow.auto.evolution.column.drop=true` short-circuits the check and the DDL commits. A failed attempt leaves a REQUESTED instant on the timeline because `client.startCommit` runs before the validation. After the DDL commits: it is its own instant, no data file is rewritten, the footer still names `v`, and `describe` shows `w variant`. **2. Reads of the renamed column with `spark.sql.variant.pushVariantIntoScan` on (the default) fail through the schema-on-read guard.** ``` org.apache.hudi.exception.HoodieException: Column 'w' is a variant requested in Spark's full-variant projection shape - by the PushVariantIntoScan rewrite (spark.sql.variant.pushVariantIntoScan) on a query, or by Hudi's own base-file reads for compaction, clustering and CDC - and the table is read with schema-on-read (hoodie.schema.on.read.enable), which cannot reconstruct variants (see issue #18285). ... ``` `ParquetSchemaEvolutionUtils.validateNoShreddedVariants`, projection-shape arm. On COW it fires from the base-file reader, on MOR from `SparkFileFormatInternalRowReaderContext` through the file group reader. Unblocked by the #18285 design: the variant column has to pass through the internal-schema pipeline as an atomic token (field-id mapping at the column level, the catalyst type spliced verbatim into the merged request, no type-change entry, no filter remap inside the synthetic struct), after which the guard's rewrite arm is retired. **3. Reads with pushdown off die in generated code.** ``` java.lang.ClassCastException: class org.apache.spark.sql.catalyst.expressions.SpecificInternalRow cannot be cast to class org.apache.spark.unsafe.types.VariantVal at org.apache.spark.sql.catalyst.expressions.BaseGenericInternalRow.getVariant(rows.scala:52) at org.apache.spark.sql.catalyst.expressions.GeneratedClass$SpecificUnsafeProjection.apply(Unknown Source) ``` Same on every leg, shredded or not. The guard's shredded-file arm resolves footer columns by the query-schema name and skips a renamed column, as its scaladoc says. The merged internal-schema request then carries `w` as `struct<metadata, value>` because the internal schema has no VARIANT arm (`InternalSchemaConverter` round-trips the sentinel record, `SparkInternalSchemaConverter.constructSparkSchemaFromInternalSchema` does not), the reader hands back a struct row where the plan's output type is variant, and the generated projection's `getVariant` cast fails. Two things unblock it: a VARIANT arm in the Spark-side converter so the merged request keeps `VariantType`, and reconstruction of shredded files (`value` plus `typed_value`) under schema-on-read, which is #18285 itself. Until then the guard could resolve renamed columns by field id so this arm fails through the guard rather than in codegen; today it is loud by accident. **4. `count(*)` fails in the vectorized reader once any internal schema is committed.** Not rename-specific: #20139, with the session-conf side effect that hides it in #20140. **5. Reads without schema-on-read after the DDL.** `w` reads null on every row, which is the schema-on-read contract (reads without it resolve by name), but the V1 catalog schema now types `w` as `struct<metadata: binary, value: binary>`: the DDL rewrote the catalog through the same converter without a VARIANT arm. Unblocked by the converter arm in 3. ## Schema-on-write (no rename primitive) **1. `ALTER TABLE ... RENAME COLUMN` without schema-on-read is refused by Spark at analysis**, with or without the drop knob: ``` org.apache.spark.sql.AnalysisException: [UNSUPPORTED_FEATURE.TABLE_OPERATION] The feature is not supported: Table `spark_catalog`.`default`.`t` does not support RENAME COLUMN. ``` `HoodieCatalog.loadTable` hands back a V1 table when schema evolution is off, and V1 tables have no rename. Nothing to unblock short of a rename primitive, and field ids only exist under schema-on-read. **2. A DataFrame write carrying `w` instead of `v`** is refused under defaults: ``` org.apache.hudi.exception.MissingSchemaFieldException: Schema validation failed due to missing field. Fields missing from incoming schema: {t_record.v} ``` With `allow.auto.evolution.column.drop=true` it commits as drop `v` plus add `w`: the table schema becomes `(id, ts, w)`, rows written before it read null for both names, and the SQL catalog still says `v` (a path write does not update it), so `w` cannot be resolved through SQL at all. A pushdown-on read of `v` over the new file then hits #20135 (fixed by #20136). That is the knob's documented drop semantics, not a rename. Conclusion: rename support is a schema-on-read feature. On schema-on-write the only work is to make the drop-plus-add outcome and the catalog divergence explicit, or refuse it. ## What supported looks like - The DDL commits under default configs (#19766). - Reads under schema-on-read return the data on both pushdown arms, COW and MOR, shredded or not (#18285 design, converter VARIANT arm, guard retired). - The catalog keeps `variant` after any schema-on-read DDL. - Reads without schema-on-read are documented as name-resolved, like any other column. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
