voonhous commented on code in PR #19691:
URL: https://github.com/apache/hudi/pull/19691#discussion_r3828603870
##########
hudi-common/src/test/resources/variant_backward_compat/README.md:
##########
@@ -39,3 +39,15 @@ The test runs on these four arguments:
COW tables generated are the same for both AVRO/SPARK. But for MOR, the log
files metadata are
different. Hence, we only need to generate test files for either 1/2, 3 and 4,
hence, 3 test
resource files.
+
+# variant_schema_on_read_cow.zip
+
+A Spark 4.1-written COW variant table carrying a COMMITTED INTERNAL SCHEMA:
one insert, then a
+schema-on-read DDL (`alter table add columns (note string)`) under
`hoodie.schema.on.read.enable`,
+then a second insert. Used by the Spark 3.x schema-on-read rejection test
(#18021). Generated by
Review Comment:
Added the exact statements to the README in d8b0c916 - create table, first
insert, the conf + alter, second insert. Recovered them from the fixture's own
commit metadata and parquet so the recorded statements match the binary exactly.
##########
hudi-spark-datasource/hudi-spark/src/test/scala/org/apache/spark/sql/hudi/dml/schema/TestVariantDataType.scala:
##########
@@ -977,6 +977,61 @@ class TestVariantDataType extends HoodieSparkSqlTestBase {
}
}
+ test("Test Spark 3.x schema-on-read reads of a variant table with a
committed internal schema") {
+ // #18021: hoodie.schema.on.read.enable resolves the schema through the
InternalSchema round
Review Comment:
Dropped in d8b0c916.
##########
hudi-spark-datasource/hudi-spark/src/test/scala/org/apache/spark/sql/hudi/dml/schema/TestVariantDataType.scala:
##########
@@ -977,6 +977,61 @@ class TestVariantDataType extends HoodieSparkSqlTestBase {
}
}
+ test("Test Spark 3.x schema-on-read reads of a variant table with a
committed internal schema") {
+ // #18021: hoodie.schema.on.read.enable resolves the schema through the
InternalSchema round
+ // trip, whose sentinel detection restores the VARIANT logical type - the
exact input
+ // HoodieSparkSchemaConverters rejects on Spark 3.x. Verified 2026-08-20:
every leg fails
+ // LOUDLY with the same actionable error as the plain auto-resolve path;
there is no silent
+ // wrong data and no obscure secondary failure. Notably that includes the
documented
+ // struct-DDL compat mode, which works WITHOUT schema-on-read (see the
backward-compat test
+ // below) but breaks once the table carries an internal schema and the
conf is on, because
+ // internal-schema resolution overrides the user's DDL. Real support is
#18285; until then
+ // Spark 3.x compat-mode readers must keep hoodie.schema.on.read.enable
off.
+ assume(HoodieSparkUtils.isSpark3, "This test verifies Spark 3.x behavior
with schema-on-read")
Review Comment:
It was meant to be CI-enforced, and the package was indeed in no Azure set.
Added org.apache.spark.sql.hudi.dml.schema to the DDL & Others wildcard set in
d8b0c916, so the spark3.5 lane now runs the pin (and the rest of the dml.schema
suites). The GHA dml lane stays spark4.2-only, where the assume cancels it as
intended.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]