Jefffrey commented on code in PR #10945:
URL: https://github.com/apache/arrow-rs/pull/10945#discussion_r4090316655
##########
arrow-select/src/take.rs:
##########
@@ -1589,6 +1637,36 @@ pub fn take_record_batch(
RecordBatch::try_new(record_batch.schema(), columns)
}
+/// Take rows by index from [`RecordBatch`], returning a new [`RecordBatch`],
without bounds
+/// checking.
+///
+/// # Safety
+///
+/// The caller must guarantee that every non-null value in `indices` is a
valid row index for
+/// `record_batch` (i.e. `index < record_batch.num_rows()`). Violating this
will cause a panic
+/// or undefined behaviour inside the inner kernels.
+///
+/// # Errors
+///
+/// Returns an [`ArrowError`] if `indices` is not an integer array type.
+pub unsafe fn take_record_batch_unchecked(
Review Comment:
im just concerned with expanding the API surface especially with unsafe
functions; for record batch we can get most of the benefits with the existing
safe API, where we only bounds check for the first column. do we get a
significant speedup by skipping the bounds check for this first column too?
for take_arrays, i suppose its unavoidable since we never enforced the
arrays to have common length, so we cant make use of a similar guarantee to
record batch 🤔
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]