andygrove opened a new issue, #6410:
URL: https://github.com/apache/datafusion-comet/issues/6410

   ### What is the problem the feature request solves?
   
   Draft PR #6404 builds Comet against DataFusion main instead of a release. 
That way, upstream changes reach Comet's CI weeks before DataFusion 56.0.0 is 
released (targeted for Oct/Nov, apache/datafusion#24461). The PR is long-lived 
and bumped weekly. It is not meant to merge before that release. It currently 
pins these versions:
   
   | Dependency     | Version                                                   
                 |
   | -------------- | 
-------------------------------------------------------------------------- |
   | DataFusion     | `4a5d580` (main, 2026-09-29)                              
                 |
   | arrow, parquet | 60.0.0                                                    
                 |
   | object_store   | 0.14.2                                                    
                 |
   | iceberg-rust   | `0d5eb12`, the head of apache/iceberg-rust#3257 (Arrow 
60, not merged yet) |
   
   This issue tracks the regressions that PR finds. The goal is to fix them 
upstream, or work around them in Comet, before the release rather than after it.
   
   ### Describe the potential solution
   
   **Regressions**
   
   Check an item off once #6404 pins a revision that has the fix.
   
   - [ ] **The native Iceberg writer puts NaN bounds in manifests.**
     - `CometIcebergWriteActionSuite` "native acceleration: NaN float/double 
manifest metrics match the JVM writer" fails 
([run](https://github.com/apache/datafusion-comet/actions/runs/36588826989/job/109492720911)).
 The native writer produces `[1,0.0,NaN,2,0.25,NaN,3,NaN,NaN]` where the JVM 
writer produces `[1,0.0,2.5,2,0.25,1.5,3,null,null]`.
     - Cause: Parquet 60 now writes NaN min/max statistics for a column chunk 
whose values are all NaN, where 59 wrote none (`get_min_max` in 
`parquet/src/column/writer/encoder.rs`). iceberg-rust's 
`MinMaxColAggregator::update` copies the Parquet statistics into lower and 
upper bounds without filtering out NaN. So a data file with an all-NaN float or 
double column gets NaN bounds, where iceberg-java writes none.
     - Impact: iceberg-java readers treat a NaN bound as unreliable, so this 
should not cause wrong results there. But the manifests no longer match the JVM 
writer's.
     - The fix belongs in iceberg-rust, ideally as part of its Arrow 60 bump 
(apache/iceberg-rust#3257). Not reported upstream yet.
   
   **Behavior changes reviewed and accepted**
   
   - Arrow 60 accepts empty entries in unsorted Variant dictionaries 
(apache/arrow-rs#10352). Comet's empty-key workaround used to reject two 
malformed dictionaries. Both now pass through unchanged, as they do in Spark.
   - A failed read from a JVM input stream now reports `Cannot get next batch 
from input stream. Error code: N. Producer error: ...`. This is because Arrow 
60 made the C Stream callbacks private, so Comet now uses the stock 
`ArrowArrayStreamReader`.
   - An invalid shuffle IPC schema message now returns a decode error instead 
of panicking (`try_fb_to_schema`).
   - DataFusion's new `floor`/`ceil` return types and `map_extract` absent-key 
results don't reach Comet, because Comet registers its own implementations.
   - DataFusion now treats a missing Parquet null count as unknown rather than 
zero. That is a correctness fix. It can reduce pruning on files written by 
parquet-rs versions before 53.1.
   
   **Upstream follow-ups**
   
   - [ ] Move the iceberg-rust pin to a main revision once 
apache/iceberg-rust#3257 merges.
   - [ ] DataFusion 56 deprecates `HashJoinExec::with_dynamic_filter_expr` with 
no replacement, and `DynamicFilterJoinExec` still needs it. File a DataFusion 
issue that describes the use case.
   - [ ] Remove the Variant workarounds that Arrow 60 makes unnecessary (#5477).
   
   **Tested revisions**
   
   | Date       | DataFusion | iceberg-rust | Suites                            
                                                 | Result                       
  |
   | ---------- | ---------- | ------------ | 
----------------------------------------------------------------------------------
 | ------------------------------ |
   | 2026-09-29 | `4a5d580`  | `0d5eb12`    | PR tier 
([run](https://github.com/apache/datafusion-comet/actions/runs/36588826989)) | 
1 failure (the NaN bounds bug) |
   
   ### Additional context
   
   Steps for triaging a red run on #6404:
   
   1. Re-run it once and compare with main's latest nightly, since some 
failures are known flakes.
   2. Run the failing test on the merge commit with the previous upstream 
revision. This tells an upstream change apart from a Comet change on main.
   3. Bisect the upstream commits between the two revisions, using a local 
`[patch]` override. DataFusion main gets 100 to 120 commits a week, so this 
takes about 7 steps.
   
   Spark's SQL tests, Iceberg's Spark tests and the non-default Spark profiles 
have not run on #6404 yet. They will run once the `run-spark-4.1-tests`, 
`run-iceberg-tests` and `run-all-spark-profiles` labels are applied.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to