Rachelint opened a new pull request, #25959: URL: https://github.com/apache/datafusion/pull/25959
## Which issue does this PR close? Part of #20773. This experimental follow-up is based on #25724; it does not close the issue. ## Rationale for this change Final aggregation with bucketed state currently builds one small `RecordBatch` per bucket, then coalesces its columns in bucket order. For a multi-column input, that repeatedly switches between columns while routing one gathered input batch. Processing all bucket ranges of a column before moving to the next column improves locality in the routing step. ## What changes are included in this PR? - Route each gathered column across the 64 buckets with Arrow's public in-progress array interface and range source API. Keep the existing coalescer's per-array copy and compaction decisions, as well as bucket spill and compaction behavior. - Pin the Arrow fork commit that exposes the public interface and range source API. This is a POC dependency until the Arrow API is upstreamed. - Add `bucket_routing_time` to aggregate metrics so the routing portion can be inspected separately. This branch contains #25724 because it is based on Jay's current head, `f9df89e2b1077deaae9be106682c9f1e4d0d61d2`. Review the final commit for this follow-up. ## What is the testing strategy for this PR? - `cargo fmt --all --check` - `cargo clippy --all-targets --all-features --offline -- -D warnings` - `cargo test -p datafusion-physical-plan final_buckets --lib --offline` (3 passed) For a benchmark bot with the standard ClickBench partitioned dataset, compare this branch against Jay's head using the same data, machine, release profile, and enabled bucket threshold. The existing `bench.sh` entry point writes comparable JSON results; no custom Criterion benchmark is needed. ```bash # Supply the same ClickBench data directory to both runs; it must contain hits_partitioned/. export DATA_DIR=/path/to/clickbench-data export DATAFUSION_EXECUTION_HASH_AGGREGATE_BUCKET_THRESHOLD=262144 DATAFUSION_DIR=/path/to/jay-25724 RESULTS_NAME=jay-25724 ./benchmarks/bench.sh run clickbench_partitioned DATAFUSION_DIR=/path/to/this-branch RESULTS_NAME=column-buckets ./benchmarks/bench.sh run clickbench_partitioned ./benchmarks/bench.sh compare jay-25724 column-buckets ``` Run `clickbench_1` the same way if desired. The bucket option defaults to off, so the environment variable is required to exercise this change. The ClickBench dataset is not available in the local checkout; the full ClickBench comparison remains for the benchmark bot. ## Are there any user-facing changes? No default behavior change. The bucketed aggregation option remains experimental and off by default. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
