andygrove commented on issue #6399:
URL: 
https://github.com/apache/datafusion-comet/issues/6399#issuecomment-5933124195

   Phase 3, scans: all 15 Parquet, Iceberg and Delta PRs have been reviewed 
against 1.0.0 and against Spark. The reproducers were then run on 1.0.0 and 
1.1.0-rc1 builds, which covers Phase 4 for this area.
   
   Three more regressions that ship in 1.1.0 are confirmed, and #6402 has the 
details:
   
   - #6504, from #5732: on an Iceberg table where a struct column, or an 
array's element struct, gained a nested field, `IS NULL` and `IS NOT NULL` on 
the column, non-outer `explode` of the array, and joins on the struct now fail. 
1.0.0 fell back to Spark for these. The root cause is older 
(apache/iceberg-rust#2617): a plain read of such a column fails natively in 
1.0.0 too.
   - #6505, from #5237: `_metadata.file_block_start` and `file_block_length` 
come out wrong when Spark splits a Parquet file into several partitions, 
because DataFusion and parquet-mr assign row groups to splits differently. 
1.0.0 fell back to Spark for metadata columns.
   - #6506, from #5681: an empty or fully pruned Parquet file with a nested 
field that the read schema can't convert now fails the query when the file is 
opened. Spark and 1.0.0 return the other files' rows. For such files that do 
have rows, the same change fixes a wrong answer.
   
   All three have a workaround in the draft release notes (#6469), checked 
against the reproducers.
   
   No regression was found in #5262 (its scan side), #5377, #5503, #5602, 
#5654, #5715, #5853, #6116, #6154 or #6219. #6154 fixes wrong rows that 1.0.0 
returned for some Iceberg transform residuals.
   
   Two findings aren't counted:
   
   - #5786 now rejects Parquet files with byte-identical duplicate field names 
in a few shapes that 1.0.0 read correctly. That's deliberate and documented, 
and only writers other than Spark produce such names.
   - #5177's TIMESTAMP_MILLIS overflow check covers a whole 8192-row batch, so 
a `LIMIT` can fail on a value Spark never reads. That needs timestamps past 
year 294,000, which Spark can't write.
   
   Why review missed these:
   
   - #5732 was argued from residual pushdown and Iceberg's post-scan filter, 
which do hold, but nobody asked what else the removed fallback was shielding, 
and no fixture evolves a nested schema.
   - #5237 reasoned that the metadata columns are constant per split. That's 
true, but which split reads a row group is a reader decision, and DataFusion 
and parquet-mr make it differently. Its tests used files too small to split.
   - #5681 moved a check that Spark makes per value to file open, and its tests 
always had rows to decode.
   - #5786's reviews measured it against `main`, which already had #5602's 
behavior, rather than against 1.0.0.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to