alamb commented on PR #10136:
URL: https://github.com/apache/arrow-rs/pull/10136#issuecomment-5821962765

   > What do you think about having the fallback walk whichever set of bits is 
smaller? The current loop runs once per set bit in mask, so at 9/10 density it 
takes about 58 iterations per word. Each iteration is 9 instructions on aarch64:
   
   In my opinion, we should merge this PR and then improve the fallback code as 
a follow on issue / PR
   
   It seems like this PR is already faster even with the somewhat basic scalar 
fallback. I am sure we can all then geek out trying to improve the performance 
of the fallback with more crazy bithacks (this is a good thing)
   
   
   I looked over comments, and it looks to me like these are the only remaining 
outstanding comments about adding some additional asserts
   - https://github.com/apache/arrow-rs/pull/10136#discussion_r4087289204 from 
me and @mbutrovich  
   - https://github.com/apache/arrow-rs/pull/10136#discussion_r4094224975 from 
@mbutrovich 
   
   For fun, I will re-run my benchmark run with the latest fixes


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to