FelixYBW commented on issue #13140: URL: https://github.com/apache/gluten/issues/13140#issuecomment-5853846262
# How does the refcount cross the C Data API today? Does Gluten use C Data to cross Java/native? **Short answer:** Gluten mostly doesn't move data across JNI — it passes handles. C Data is used only when Java code needs to read or produce Arrow data. Gluten's own reference counts never cross the boundary; C Data works by moving ownership once, and Arrow's per-buffer counts and Velox's shared pointers keep the memory alive. --- ## How Java and native connect today ### 1. Handles — no data crossing (normal case between Velox operators) - A native batch lives in the C++ `ObjectStore` as a `shared_ptr<ColumnarBatch>`. Java only holds the handle, inside `IndicatorVector`. - Gluten's Java-side count controls when Java calls `ColumnarBatchJniWrapper.close(handle)`, which runs `ObjectStore::release` and drops that `shared_ptr`. - Row conversion (`VeloxColumnarToRow` / `RowToVeloxColumnar`) and shuffle also cross as serialized bytes, not C Data. ### 2. C Data — only for conversions between Java Arrow and native These are `LoadArrowDataExec` / `OffloadArrowDataExec`, used by the Python exec, partial project/generate, and the Java Arrow paths. #### Native → Java (`ColumnarBatches.load`) 1. Java calls `ColumnarBatchJniWrapper.exportToArrow(handle, cSchema, cArray)`. 2. In C++ (`JniWrapper.cc`), `VeloxColumnarBatch::exportArrowArray()` calls `velox::exportToArrow(rowVector_, ...)`. The exported `ArrowArray` keeps its own references to the Velox buffers through its private data and release callback. The struct is then handed over with `ArrowArrayMove`. 3. Back in Java, `ArrowAbiUtil` calls `Data.importVectorSchemaRoot` / `importIntoVectorSchemaRoot`. Arrow Java wraps the foreign buffers as `ArrowBuf`s without copying. When the last of them is released, Arrow Java calls the C release callback once, and Velox drops its references. 4. `load` then closes the input native batch (it follows the count, as described earlier). That's safe because the exported array keeps the Velox buffers alive independently. #### Java → native (`ColumnarBatches.offload`) 1. `ArrowAbiUtil.exportFromSparkColumnarBatch` calls `Data.exportVectorSchemaRoot`. Arrow Java's exporter retains each `ArrowBuf` (its own count +1) and sets a release callback that drops those references. 2. `createWithArrowArray` moves the struct into an `ArrowCStructColumnarBatch` (`ArrowArrayMove`). Its destructor calls `ArrowArrayRelease`. 3. `ArrowColumnarToVeloxColumnarExec` / `VeloxColumnarBatches.toVeloxBatch` converts it to Velox with `velox::importFromArrowAsOwner(...)` (`VeloxColumnarBatch.cc:101`). Velox takes over the array, and the release callback runs when the Velox vectors are freed. 4. The Java batch can then be closed normally. The exported buffers survive through Arrow's per-buffer counts. --- ## Three layers of lifetime | Layer | Mechanism | Crosses JNI? | |---|---|---| | Gluten Java objects | `IndicatorVector.refCnt`, `ArrowWritableColumnVector.refCnt` | No. They only decide when Java calls `close()` / `ObjectStore::release` | | Arrow Java buffers | `ArrowBuf` reference counts | Through C Data: export retains the buffers; import creates buffers whose last release calls the C release callback | | Native memory | `shared_ptr` in `ObjectStore`, Velox buffer reference counts, references held by the exported array | Through C Data's release callback | --- > **C Data has no reference count field.** Its model is a single move of ownership: whoever holds the struct must call `release` exactly once. Gluten follows that with `ArrowArrayMove` on both sides. --- ## What this means for the switch - The C Data path doesn't change. Import and export work the same with `ArrowColumnVector`: `Data.importVectorSchemaRoot`, then wrap the vectors. For export: `getValueVector` → `VectorSchemaRoot.of` → `Data.exportVectorSchemaRoot`. - Lifetime across the boundary is already handled *below* Gluten's count, by Arrow's buffer counts and the one-time release callback. That's why the Java Arrow vector-level count can go: once data is exported, the Java batch can close whenever its producer is done. - The native-side count is separate. The `IndicatorVector` count still governs native batches that several Java holders keep, such as limit, tail, and broadcast. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
