zhuqi-lucas opened a new pull request, #26104: URL: https://github.com/apache/datafusion/pull/26104
## Which issue does this PR close? - Closes #26103 ## Rationale for this change `ListingTable` schema inference over Parquet merges the per-file schemas with `Schema::try_merge`, which widens nullability only for fields that appear in more than one schema. A column that is `required` in some files and absent from the others therefore comes out `NOT NULL`, but reading the files without it fills it with nulls, and the scan fails: ``` Arrow error: Invalid argument error: Column 'c' is declared as non-nullable but contains null values ``` Pure SQL reproducer (on `main`): ```sql COPY (SELECT 1 AS id, 10 AS c) TO '/tmp/evo/a.parquet'; COPY (SELECT 2 AS id) TO '/tmp/evo/b.parquet'; CREATE EXTERNAL TABLE t STORED AS PARQUET LOCATION '/tmp/evo/'; SELECT * FROM t; -- fails; DESCRIBE shows c Int64 NO ``` Before DataFusion 52 this was masked by `SchemaAdapter::map_batch`, which rebuilt batches under the table schema without validating nullability (removed in #18998). The strict check is right; the inferred schema is what is wrong. ## What changes are included in this PR? In `ParquetFormat::infer_schema`, count in how many files each top-level field appears and, after the merge, mark every field that is missing from at least one file as nullable. Fields present in every file keep whatever `try_merge` produced, so tables whose files all share a schema are unchanged. ## Are these changes tested? - New `schema_evolution.slt` case: two `COPY TO` files, `CREATE EXTERNAL TABLE` without a schema, `DESCRIBE` shows the partial column as nullable, `SELECT` returns the null. Fails on `main` with the error above. - New `schema_merge_marks_partially_present_columns_nullable` in `core/tests/parquet/schema.rs`: three files, column in a different position, an already-nullable partial column, and both `skip_metadata` paths. ## Are there any user-facing changes? An inferred Parquet table schema now reports a column as nullable when some files do not contain it. Tables where every file has every column are unaffected. Explicitly declared schemas are not touched. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
