bharadwaj-pendyala opened a new pull request, #11256:
URL: https://github.com/apache/arrow-rs/pull/11256

   # Which issue does this PR close?
   
   - Part of #11200.
   
   # Rationale for this change
   
   #11200 is about `FilterPredicate::filter_nulls` building and then throwing 
away a validity bitmap when the filter keeps no null rows. None of the existing 
`filter_kernels` cases hit that: every nullable input is 50% random nulls 
filtered by a mask that ignores them, so the output always has nulls.
   
   CONTRIBUTING asks for new benchmarks in their own PR so the bench runner can 
compare against them, so these come ahead of the kernel change.
   
   # What changes are included in this PR?
   
   Five cases in `arrow/benches/filter_kernels.rs`, next to the existing `i32 w 
NULLs` ones:
   
   - `filter context i32 w NULLs, only valid` at kept 1/4, 1023/2048 and 
1/2048. Each mask is the existing 1/2, dense or sparse mask ANDed with 
`is_not_null` of the data, so it never selects a null.
   - `filter context i32 w NULLs at end` at kept 1/2 and 1023/1024, over an 
array whose only nulls are the last 64 rows. These are the adverse case the 
issue asks for: any check for a selected null has to scan to the last word 
before it finds one.
   
   # Are these changes tested?
   
   Benchmarks only. `cargo clippy -p arrow --bench filter_kernels --features 
test_utils -- -D warnings` is clean and all five run.
   
   Baseline on main at `fa337f8`, M1 laptop, criterion median:
   
   ```
   i32 w NULLs, only valid (kept 1/4)                                ~120 µs
   i32 w NULLs, only valid high selectivity (kept 1023/2048)         ~221 µs
   i32 w NULLs, only valid low selectivity (kept 1/2048)             ~1.10 µs
   i32 w NULLs at end (kept 1/2)                                     ~221 µs
   i32 w NULLs at end high selectivity (kept 1023/1024)              ~42 µs
   ```
   
   This machine is noisy: the same binary moved between 73 µs and 135 µs on the 
first case across runs, so read these as rough.
   
   # Are there any user-facing changes?
   
   No.
   
   # AI usage
   
   Claude wrote the benchmarks and Codex reviewed them adversarially; its first 
pass is what asked for the late-null cases. I ran every number above myself.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to