devanbenz commented on PR #10136: URL: https://github.com/apache/arrow-rs/pull/10136#issuecomment-5822684602
> > What do you think about having the fallback walk whichever set of bits is smaller? The current loop runs once per set bit in mask, so at 9/10 density it takes about 58 iterations per word. Each iteration is 9 instructions on aarch64: > > In my opinion, we should merge this PR and then improve the fallback code as a follow on issue / PR > > It seems like this PR is already faster even with the somewhat basic scalar fallback. I am sure we can all then geek out trying to improve the performance of the fallback with more crazy bithacks (this is a good thing) > > I looked over comments, and it looks to me like these are the only remaining outstanding comments about adding some additional asserts > > * [feat(perf): Improve filter performance with per word bit filtering (and `BMI` when supported) #10136 (comment)](https://github.com/apache/arrow-rs/pull/10136#discussion_r4087289204) from me and @mbutrovich > > * [feat(perf): Improve filter performance with per word bit filtering (and `BMI` when supported) #10136 (comment)](https://github.com/apache/arrow-rs/pull/10136#discussion_r4094224975) from @mbutrovich > > > For fun, I will re-run my benchmark run with the latest fixes I've gone ahead and added the suggested `debug_asserts`. I've also added a new benchmark for FSB which includes nulls to stress the code path in `filter_bits`. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
