alamb commented on PR #10136: URL: https://github.com/apache/arrow-rs/pull/10136#issuecomment-5821962765
> What do you think about having the fallback walk whichever set of bits is smaller? The current loop runs once per set bit in mask, so at 9/10 density it takes about 58 iterations per word. Each iteration is 9 instructions on aarch64: In my opinion, we should merge this PR and then improve the fallback code as a follow on issue / PR It seems like this PR is already faster even with the somewhat basic scalar fallback. I am sure we can all then geek out trying to improve the performance of the fallback with more crazy bithacks (this is a good thing) I looked over comments, and it looks to me like these are the only remaining outstanding comments about adding some additional asserts - https://github.com/apache/arrow-rs/pull/10136#discussion_r4087289204 from me and @mbutrovich - https://github.com/apache/arrow-rs/pull/10136#discussion_r4094224975 from @mbutrovich For fun, I will re-run my benchmark run with the latest fixes -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
