limadog9 opened a new pull request, #25818: URL: https://github.com/apache/datafusion/pull/25818
## Which issue does this PR close? Closes #25817. ## Rationale for this change When sorted Parquet files use different timestamp units from the declared table schema, `ORDER BY` can return files in listing order and timestamp predicates lose file-level pruning. Failed statistics updates were leaving exact NULL bounds that incorrectly validated the declared ordering. ## What changes are included in this PR? - Convert timestamp bounds to the table's unit before accumulation, using checked casts when the timezone is unchanged. - Discard min/max bounds when statistics summarization fails, while retaining null counts and byte sizes. - Mark distinct counts inexact when type conversion can merge values, avoiding incorrect `COUNT(DISTINCT)` substitutions. - Add SQL regression coverage for file ordering, LIMIT, aggregates, file pruning, and negative timestamp downcasts; add metadata tests for overflow, incompatible types, distinct counts, and mixed bound exactness. ## What is the testing strategy for this PR? Completed locally on Windows: - `cargo fmt --all` and `git diff --check`. - `cargo test --profile ci -p datafusion-datasource-parquet --lib`: 268 tests passed. - `parquet_sorted_statistics.slt`: passed using the built sqllogictest binary, including the assertion that timestamp file-range pruning matches 1 of 3 files. This PR is a draft while broader validation is pending: - The initial `cargo clippy --all-targets --all-features -- -D warnings` run encountered a pre-existing Windows-only `clippy::unnecessary_semicolon` warning in `datafusion/common/src/rounding.rs:257`. This PR leaves that file unchanged. - The broader workspace test and benchmark builds were stopped before completion to publish this draft. The benchmark comparison with `main` and the regression run with the fix removed remain pending. - `dev/rust_lint.sh` stopped at its Python/PyYAML prerequisite check. ## Are there any user-facing changes? Queries over Parquet files whose timestamp units differ from the table schema return correctly ordered rows and can use timestamp statistics for file pruning. No public API changes. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
