Rich-T-kid commented on code in PR #10945:
URL: https://github.com/apache/arrow-rs/pull/10945#discussion_r4094162085
##########
arrow-select/src/take.rs:
##########
@@ -1589,6 +1637,36 @@ pub fn take_record_batch(
RecordBatch::try_new(record_batch.schema(), columns)
}
+/// Take rows by index from [`RecordBatch`], returning a new [`RecordBatch`],
without bounds
+/// checking.
+///
+/// # Safety
+///
+/// The caller must guarantee that every non-null value in `indices` is a
valid row index for
+/// `record_batch` (i.e. `index < record_batch.num_rows()`). Violating this
will cause a panic
+/// or undefined behaviour inside the inner kernels.
+///
+/// # Errors
+///
+/// Returns an [`ArrowError`] if `indices` is not an integer array type.
+pub unsafe fn take_record_batch_unchecked(
Review Comment:
> for record batch we can get most of the benefits with the existing safe
API, where we only bounds check for the first column.
This is true
> do we get a significant speedup by skipping the bounds check for this
first column too?
no id expect it to be proportional to the number of columns in the record
batch, maybe besides some branch mis-prediction but this is quite minor
Will update 👍
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]