andygrove opened a new pull request, #6540:
URL: https://github.com/apache/datafusion-comet/pull/6540

   ## Which issue does this PR close?
   
   Closes #6506 on `branch-1.1`, for 1.1.0-rc2.
   
   ## Rationale for this change
   
   This is the `branch-1.1` backport of #6515. Since #5681, the native Parquet 
scan checks every nested field's type conversion when it opens a file, where 
Spark checks it only when it decodes a row group. So a file with nothing to 
decode, empty or with every row group pruned, whose nested field has a type the 
read schema can't convert, fails the query on 1.1.0-rc1, where Spark and 1.0.0 
return the other files' rows. #6515 has the details.
   
   `branch-1.0` doesn't need it. It doesn't have #5681, and the top-level part 
of #6515 isn't a regression fix, since 1.0.0 already rejected those pairs when 
it opened the file.
   
   ## What changes are included in this PR?
   
   A clean cherry-pick (`-x`) of #6515's commit, with no adaptations. 
Conversion errors that Spark raises while decoding are now raised when a row 
group is decoded, and only the shape mismatches that Spark also rejects when it 
opens the file still fail at open. The scan compatibility guide documents the 
one gap left, with `spark.comet.parquet.rowFilterPushdown.enabled=true`, which 
is off by default.
   
   ## How are these changes tested?
   
   On this branch:
   
   - `cargo test -p datafusion-comet --lib`: 539 passed, including the four new 
`schema_adapter` tests. `schema_adapter.rs` and 
`eager_page_index_reader_factory.rs` are byte-identical to `main`'s, both 
before and after #6515, and on `main` all four tests fail without the change.
   - Workspace clippy (`--all-targets -D warnings`) and rustfmt are clean.
   - `ParquetReadV1Suite` on the default Spark 4.1 profile: 80 passed and 1 
ignored, including the new "native scan reads files with nothing to decode 
whose types Spark rejects".
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to