Rachelint opened a new pull request, #25959:
URL: https://github.com/apache/datafusion/pull/25959

   ## Which issue does this PR close?
   
   Part of #20773. This experimental follow-up is based on #25724; it does not 
close the issue.
   
   ## Rationale for this change
   
   Final aggregation with bucketed state currently builds one small 
`RecordBatch` per bucket, then coalesces its columns in bucket order. For a 
multi-column input, that repeatedly switches between columns while routing one 
gathered input batch. Processing all bucket ranges of a column before moving to 
the next column improves locality in the routing step.
   
   ## What changes are included in this PR?
   
   - Route each gathered column across the 64 buckets with Arrow's public 
in-progress array interface and range source API. Keep the existing coalescer's 
per-array copy and compaction decisions, as well as bucket spill and compaction 
behavior.
   - Pin the Arrow fork commit that exposes the public interface and range 
source API. This is a POC dependency until the Arrow API is upstreamed.
   - Add `bucket_routing_time` to aggregate metrics so the routing portion can 
be inspected separately.
   
   This branch contains #25724 because it is based on Jay's current head, 
`f9df89e2b1077deaae9be106682c9f1e4d0d61d2`. Review the final commit for this 
follow-up.
   
   ## What is the testing strategy for this PR?
   
   - `cargo fmt --all --check`
   - `cargo clippy --all-targets --all-features --offline -- -D warnings`
   - `cargo test -p datafusion-physical-plan final_buckets --lib --offline` (3 
passed)
   
   For a benchmark bot with the standard ClickBench partitioned dataset, 
compare this branch against Jay's head using the same data, machine, release 
profile, and enabled bucket threshold. The existing `bench.sh` entry point 
writes comparable JSON results; no custom Criterion benchmark is needed.
   
   ```bash
   # Supply the same ClickBench data directory to both runs; it must contain 
hits_partitioned/.
   export DATA_DIR=/path/to/clickbench-data
   export DATAFUSION_EXECUTION_HASH_AGGREGATE_BUCKET_THRESHOLD=262144
   
   DATAFUSION_DIR=/path/to/jay-25724 RESULTS_NAME=jay-25724 
./benchmarks/bench.sh run clickbench_partitioned
   DATAFUSION_DIR=/path/to/this-branch RESULTS_NAME=column-buckets 
./benchmarks/bench.sh run clickbench_partitioned
   ./benchmarks/bench.sh compare jay-25724 column-buckets
   ```
   
   Run `clickbench_1` the same way if desired. The bucket option defaults to 
off, so the environment variable is required to exercise this change. The 
ClickBench dataset is not available in the local checkout; the full ClickBench 
comparison remains for the benchmark bot.
   
   ## Are there any user-facing changes?
   
   No default behavior change. The bucketed aggregation option remains 
experimental and off by default.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to