limadog9 opened a new pull request, #25818:
URL: https://github.com/apache/datafusion/pull/25818

   ## Which issue does this PR close?
   
   Closes #25817.
   
   ## Rationale for this change
   
   When sorted Parquet files use different timestamp units from the declared 
table schema, `ORDER BY` can return files in listing order and timestamp 
predicates lose file-level pruning. Failed statistics updates were leaving 
exact NULL bounds that incorrectly validated the declared ordering.
   
   ## What changes are included in this PR?
   
   - Convert timestamp bounds to the table's unit before accumulation, using 
checked casts when the timezone is unchanged.
   - Discard min/max bounds when statistics summarization fails, while 
retaining null counts and byte sizes.
   - Mark distinct counts inexact when type conversion can merge values, 
avoiding incorrect `COUNT(DISTINCT)` substitutions.
   - Add SQL regression coverage for file ordering, LIMIT, aggregates, file 
pruning, and negative timestamp downcasts; add metadata tests for overflow, 
incompatible types, distinct counts, and mixed bound exactness.
   
   ## What is the testing strategy for this PR?
   
   Completed locally on Windows:
   
   - `cargo fmt --all` and `git diff --check`.
   - `cargo test --profile ci -p datafusion-datasource-parquet --lib`: 268 
tests passed.
   - `parquet_sorted_statistics.slt`: passed using the built sqllogictest 
binary, including the assertion that timestamp file-range pruning matches 1 of 
3 files.
   
   This PR is a draft while broader validation is pending:
   
   - The initial `cargo clippy --all-targets --all-features -- -D warnings` run 
encountered a pre-existing Windows-only `clippy::unnecessary_semicolon` warning 
in `datafusion/common/src/rounding.rs:257`. This PR leaves that file unchanged.
   - The broader workspace test and benchmark builds were stopped before 
completion to publish this draft. The benchmark comparison with `main` and the 
regression run with the fix removed remain pending.
   - `dev/rust_lint.sh` stopped at its Python/PyYAML prerequisite check.
   
   ## Are there any user-facing changes?
   
   Queries over Parquet files whose timestamp units differ from the table 
schema return correctly ordered rows and can use timestamp statistics for file 
pruning. No public API changes.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to