andygrove opened a new issue, #6410:
URL: https://github.com/apache/datafusion-comet/issues/6410
### What is the problem the feature request solves?
Draft PR #6404 builds Comet against DataFusion main instead of a release.
That way, upstream changes reach Comet's CI weeks before DataFusion 56.0.0 is
released (targeted for Oct/Nov, apache/datafusion#24461). The PR is long-lived
and bumped weekly. It is not meant to merge before that release. It currently
pins these versions:
| Dependency | Version
|
| -------------- |
-------------------------------------------------------------------------- |
| DataFusion | `4a5d580` (main, 2026-09-29)
|
| arrow, parquet | 60.0.0
|
| object_store | 0.14.2
|
| iceberg-rust | `0d5eb12`, the head of apache/iceberg-rust#3257 (Arrow
60, not merged yet) |
This issue tracks the regressions that PR finds. The goal is to fix them
upstream, or work around them in Comet, before the release rather than after it.
### Describe the potential solution
**Regressions**
Check an item off once #6404 pins a revision that has the fix.
- [ ] **The native Iceberg writer puts NaN bounds in manifests.**
- `CometIcebergWriteActionSuite` "native acceleration: NaN float/double
manifest metrics match the JVM writer" fails
([run](https://github.com/apache/datafusion-comet/actions/runs/36588826989/job/109492720911)).
The native writer produces `[1,0.0,NaN,2,0.25,NaN,3,NaN,NaN]` where the JVM
writer produces `[1,0.0,2.5,2,0.25,1.5,3,null,null]`.
- Cause: Parquet 60 now writes NaN min/max statistics for a column chunk
whose values are all NaN, where 59 wrote none (`get_min_max` in
`parquet/src/column/writer/encoder.rs`). iceberg-rust's
`MinMaxColAggregator::update` copies the Parquet statistics into lower and
upper bounds without filtering out NaN. So a data file with an all-NaN float or
double column gets NaN bounds, where iceberg-java writes none.
- Impact: iceberg-java readers treat a NaN bound as unreliable, so this
should not cause wrong results there. But the manifests no longer match the JVM
writer's.
- The fix belongs in iceberg-rust, ideally as part of its Arrow 60 bump
(apache/iceberg-rust#3257). Not reported upstream yet.
**Behavior changes reviewed and accepted**
- Arrow 60 accepts empty entries in unsorted Variant dictionaries
(apache/arrow-rs#10352). Comet's empty-key workaround used to reject two
malformed dictionaries. Both now pass through unchanged, as they do in Spark.
- A failed read from a JVM input stream now reports `Cannot get next batch
from input stream. Error code: N. Producer error: ...`. This is because Arrow
60 made the C Stream callbacks private, so Comet now uses the stock
`ArrowArrayStreamReader`.
- An invalid shuffle IPC schema message now returns a decode error instead
of panicking (`try_fb_to_schema`).
- DataFusion's new `floor`/`ceil` return types and `map_extract` absent-key
results don't reach Comet, because Comet registers its own implementations.
- DataFusion now treats a missing Parquet null count as unknown rather than
zero. That is a correctness fix. It can reduce pruning on files written by
parquet-rs versions before 53.1.
**Upstream follow-ups**
- [ ] Move the iceberg-rust pin to a main revision once
apache/iceberg-rust#3257 merges.
- [ ] DataFusion 56 deprecates `HashJoinExec::with_dynamic_filter_expr` with
no replacement, and `DynamicFilterJoinExec` still needs it. File a DataFusion
issue that describes the use case.
- [ ] Remove the Variant workarounds that Arrow 60 makes unnecessary (#5477).
**Tested revisions**
| Date | DataFusion | iceberg-rust | Suites
| Result
|
| ---------- | ---------- | ------------ |
----------------------------------------------------------------------------------
| ------------------------------ |
| 2026-09-29 | `4a5d580` | `0d5eb12` | PR tier
([run](https://github.com/apache/datafusion-comet/actions/runs/36588826989)) |
1 failure (the NaN bounds bug) |
### Additional context
Steps for triaging a red run on #6404:
1. Re-run it once and compare with main's latest nightly, since some
failures are known flakes.
2. Run the failing test on the merge commit with the previous upstream
revision. This tells an upstream change apart from a Comet change on main.
3. Bisect the upstream commits between the two revisions, using a local
`[patch]` override. DataFusion main gets 100 to 120 commits a week, so this
takes about 7 steps.
Spark's SQL tests, Iceberg's Spark tests and the non-default Spark profiles
have not run on #6404 yet. They will run once the `run-spark-4.1-tests`,
`run-iceberg-tests` and `run-all-spark-profiles` labels are applied.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]