viirya opened a new pull request, #6534:
URL: https://github.com/apache/datafusion-comet/pull/6534

   ## Which issue does this PR close?
   
   Part of #6434.
   
   ## Rationale for this change
   
   #6435 separated the error classification at the native boundary and the core 
logic of a few JNI
   entry points. This continues with the remaining entry points that do not 
call back into the JVM,
   so their native logic can be unit-tested without a JVM.
   
   ## What changes are included in this PR?
   
   Each export keeps converting its JNI arguments and return value, and calls a 
core function with no
   JNI types in its signature:
   
   - `columnarToRowConvert`: `columnar_to_row_convert(ctx, &[i64], &[i64], 
num_rows)` imports the
     Arrow C Data columns and converts them. The export only pins the address 
arrays and builds
     `NativeColumnarToRowInfo`.
   - `writeSortedFileNative`: `write_sorted_file(&[i64], &[i32], ...)` takes 
the row addresses and
     sizes as slices and handles the codec name and the `i64::MIN` checksum 
convention, returning
     `[bytes written, checksum, encode nanos]`.
   - `sortRowPartitionsNative`: `sort_row_partitions(&mut [i64])`.
   - `getShufflePartitionOffsets`: `shuffle_partition_offsets(Option<&dyn 
ExecutionPlan>)` takes the
     root plan instead of the execution context.
   - `NativeBase.init`: `init_logging` and `logger_config` set up logging. The 
export keeps capturing
     the `JavaVM` and installing the panic hook.
   - `NativeBase.getTzdataVersion`: `tzdata_version()`.
   
   `columnarToRowInit`/`columnarToRowClose`, `NativeBase.release` and the 
tracing entry points already
   only call JNI-free functions and are unchanged. Exception classes and 
messages are unchanged.
   
   ## How are these changes tested?
   
   - New Rust unit tests without a JVM for `columnar_to_row_convert` (compared 
with converting the same
     arrays directly), `write_sorted_file` (bytes written match the file, 
checksum present only when
     enabled and stable), `sort_row_partitions` (stable sort by partition id),
     `shuffle_partition_offsets` (errors before execution and for a plan that 
is not a shuffle
     write), `logger_config` and `tzdata_version`.
   - Ran `CometNativeColumnarToRowSuite`, `CometShuffleSuite`, 
`CometNativeShuffleSuite` and
     `CometTemporalExpressionSuite` locally.
   
   This pull request and its description were written by Isaac.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to