andygrove opened a new pull request, #6540: URL: https://github.com/apache/datafusion-comet/pull/6540
## Which issue does this PR close? Closes #6506 on `branch-1.1`, for 1.1.0-rc2. ## Rationale for this change This is the `branch-1.1` backport of #6515. Since #5681, the native Parquet scan checks every nested field's type conversion when it opens a file, where Spark checks it only when it decodes a row group. So a file with nothing to decode, empty or with every row group pruned, whose nested field has a type the read schema can't convert, fails the query on 1.1.0-rc1, where Spark and 1.0.0 return the other files' rows. #6515 has the details. `branch-1.0` doesn't need it. It doesn't have #5681, and the top-level part of #6515 isn't a regression fix, since 1.0.0 already rejected those pairs when it opened the file. ## What changes are included in this PR? A clean cherry-pick (`-x`) of #6515's commit, with no adaptations. Conversion errors that Spark raises while decoding are now raised when a row group is decoded, and only the shape mismatches that Spark also rejects when it opens the file still fail at open. The scan compatibility guide documents the one gap left, with `spark.comet.parquet.rowFilterPushdown.enabled=true`, which is off by default. ## How are these changes tested? On this branch: - `cargo test -p datafusion-comet --lib`: 539 passed, including the four new `schema_adapter` tests. `schema_adapter.rs` and `eager_page_index_reader_factory.rs` are byte-identical to `main`'s, both before and after #6515, and on `main` all four tests fail without the change. - Workspace clippy (`--all-targets -D warnings`) and rustfmt are clean. - `ParquetReadV1Suite` on the default Spark 4.1 profile: 80 passed and 1 ignored, including the new "native scan reads files with nothing to decode whose types Spark rejects". -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
