peterxcli opened a new issue, #11260: URL: https://github.com/apache/arrow-rs/issues/11260
### Is your feature request related to a problem or challenge? Comet currently calls `unshred_variant`, then decodes and rebuilds its output to match Spark's Variant bytes (apache/datafusion-comet#5978). We need to reuse Arrow's typed decoders and recursive traversal while writing the final representation directly. The public API materializes a VariantArray using Arrow's writer. The reusable [`UnshredVariantRowBuilder`](https://github.com/apache/arrow-rs/blob/76829d5ce4562616a69ea8b2696e96160514482f/parquet-variant-compute/src/unshred_variant.rs#L160-L207) is private. Making it public alone would not provide enough control: typed and residual scalars share `append_value`, and nested output uses concrete object/list builders. ### Describe the solution you'd like Expose a reusable row decoder in `parquet-variant-compute`, constructed once per input array, with a fallible row visitor or equivalent borrowed interface. For example, `decode_row(index, visitor)`; the exact API is open for discussion. The caller needs to: - Receive typed scalars separately from residual bytes and their source metadata. Spark narrows typed integers and canonicalizes typed NaNs, while preserving those encodings in residual scalars. - Handle nested object fields and list elements through the same interface, preserving typed schema order and original residual field order. - Distinguish a null parent, a missing shredding state and explicit Variant null, with access to state validity before residual decoding. Honor parent masking and sliced list offsets. - Build output metadata independently of input metadata. Residual containers must expose their names/children so the caller can remap field IDs instead of copying containers with stale IDs. Keep `unshred_variant` as the existing canonical entry point, ideally backed by this decoder and its default writer. Spark-specific encoding and permissive-reading policies remain explicit consumer choices. An external-crate test should demonstrate direct output for partially and fully shredded nested rows, without an intermediate encoded VariantArray. Cover typed/residual provenance, differing dictionaries, nulls, sliced lists and fallible malformed-input handling; verify the existing entry point retains its behavior. A focused benchmark should report allocation and runtime differences. ### Describe alternatives you've considered Unshred then re-encode, as Comet does today, or duplicate Arrow's typed decoders downstream. Both add avoidable work or maintenance. #10620 addresses returning a materialized composite Variant from `try_value`; this request exposes traversal before materialization and could share its decoder. ### Additional context apache/datafusion-comet#5978 tracks the full Spark-compatible reconstruction and performance target. This API is one prerequisite; legacy input handling and other compatibility policies still need separate integration. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
