zhuqi-lucas opened a new pull request, #26104:
URL: https://github.com/apache/datafusion/pull/26104

   ## Which issue does this PR close?
   
   - Closes #26103
   
   ## Rationale for this change
   
   `ListingTable` schema inference over Parquet merges the per-file schemas 
with `Schema::try_merge`, which widens nullability only for fields that appear 
in more than one schema. A column that is `required` in some files and absent 
from the others therefore comes out `NOT NULL`, but reading the files without 
it fills it with nulls, and the scan fails:
   
   ```
   Arrow error: Invalid argument error: Column 'c' is declared as non-nullable 
but contains null values
   ```
   
   Pure SQL reproducer (on `main`):
   
   ```sql
   COPY (SELECT 1 AS id, 10 AS c) TO '/tmp/evo/a.parquet';
   COPY (SELECT 2 AS id)          TO '/tmp/evo/b.parquet';
   CREATE EXTERNAL TABLE t STORED AS PARQUET LOCATION '/tmp/evo/';
   SELECT * FROM t;  -- fails; DESCRIBE shows c Int64 NO
   ```
   
   Before DataFusion 52 this was masked by `SchemaAdapter::map_batch`, which 
rebuilt batches under the table schema without validating nullability (removed 
in #18998). The strict check is right; the inferred schema is what is wrong.
   
   ## What changes are included in this PR?
   
   In `ParquetFormat::infer_schema`, count in how many files each top-level 
field appears and, after the merge, mark every field that is missing from at 
least one file as nullable. Fields present in every file keep whatever 
`try_merge` produced, so tables whose files all share a schema are unchanged.
   
   ## Are these changes tested?
   
   - New `schema_evolution.slt` case: two `COPY TO` files, `CREATE EXTERNAL 
TABLE` without a schema, `DESCRIBE` shows the partial column as nullable, 
`SELECT` returns the null. Fails on `main` with the error above.
   - New `schema_merge_marks_partially_present_columns_nullable` in 
`core/tests/parquet/schema.rs`: three files, column in a different position, an 
already-nullable partial column, and both `skip_metadata` paths.
   
   ## Are there any user-facing changes?
   
   An inferred Parquet table schema now reports a column as nullable when some 
files do not contain it. Tables where every file has every column are 
unaffected. Explicitly declared schemas are not touched.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to