alamb commented on code in PR #10136:
URL: https://github.com/apache/arrow-rs/pull/10136#discussion_r4097497290


##########
arrow-buffer/src/util/bit_util.rs:
##########
@@ -19,6 +19,68 @@
 
 use crate::bit_chunk_iterator::BitChunks;
 
+/// Parallel bit extract: for each set bit in `mask`, extract the
+/// corresponding bit from `value` and pack them contiguously into the low
+/// bits of the return value.
+///
+/// Equivalent to the x86 BMI2 `PEXT` instruction. When compiled with the
+/// `bmi2` target feature enabled (for example `-C target-cpu=x86-64-v3`)
+/// this lowers to the hardware `pext` instruction; otherwise it falls back
+/// to a portable scalar loop.
+///
+/// # Functional Example
+///
+/// Using 8 bits for brevity (the function operates on all 64). Each
+/// set bit in `mask` selects the bit at the same position in `value`; the
+/// selected bits are then shifted down so they are contiguous in the low
+/// bits of the result, in their original order:
+///
+/// ```text
+/// bit:     7 6 5 4 3 2 1 0

Review Comment:
   Apparently there is no CI coverage for this yet (the existing release job 
doesn't use all the fancy architectural features). I made a PR to propose doing 
this
   - https://github.com/apache/arrow-rs/pull/11192



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to