devanbenz opened a new pull request, #11220:
URL: https://github.com/apache/arrow-rs/pull/11220

   A handful of benchmarks were set up with inputs that did not match what they 
claim to measure. This PR fixes the inputs so the numbers actually reflect the 
intended workload.
   
   - `buffer_create`: `MutableBuffer iter bitset` allocated 4KiB per buffer 
(from the outer vec length) and set the length to `datum.len()` bytes instead 
of the bitmap size. Now sized to `datum.len().div_ceil(8)`.
   - `interleave_kernels`: `dict_distinct` always used 100 indices regardless 
of `len`, so 1024 and 2048 were measuring the same thing.
   - `lexsort`: `Optional50CharString` was generated without nulls and one case 
was duplicated.
   - `take_kernels`: `take list i32 null indices 1024` used lists of 202 
elements instead of 20 like the rest of the list benches.
   - `arrow_reader_row_filter`: `UnselectiveUnclustered` was missing the `NOT`, 
so it was measuring the same ~1% selective filter as `SelectiveUnclustered` 
instead of ~99%.
   
   Benchmarks from my machine (i7-12700K), main vs this branch:
   
   | Bench | main | fixed | change |
   |---|---|---|---|
   | `MutableBuffer iter bitset` | 39.3 ms | 4.38 ms | -89% |
   | `interleave dict_distinct 1024` | 1.43 µs | 4.73 µs | +234% |
   | `interleave dict_distinct 2048` | 1.33 µs | 2.32 µs | +74% |
   | `take list i32 null indices 1024` | 6.29 µs | 2.76 µs | -56% |
   | `float64 <= 99.0` row filter (non-limit) | 4.31–5.17 ms | 6.68–8.31 ms | 
+53% to +61% |
   | `lexsort` cases with `str_opt(50)` | | | -6% to +20% |
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to