TaiSakuma opened a new issue, #50623:
URL: https://github.com/apache/arrow/issues/50623

   ### Describe the bug, including details regarding any error messages, 
version, and platform
   
   Serializing a table whose column is a registered `ExtensionType` with union 
storage through Arrow IPC (file or stream format) produces bytes in which the 
union's buffers are missing. Reading them back yields a table that fails 
`validate(full=True)` with `Buffer #1 too small`, and accessing the values 
(`to_pylist`) crashes the process with a segmentation fault. The table 
validates before writing, and the same union columns without the extension 
wrapper roundtrip correctly.
   
   ```python
   import io
   import pyarrow as pa
   import pyarrow.ipc
   
   
   class MyExt(pa.ExtensionType):
       def __init__(self, storage_type):
           super().__init__(storage_type, "my.ext")
   
       def __arrow_ext_serialize__(self):
           return b""
   
       @classmethod
       def __arrow_ext_deserialize__(cls, storage_type, serialized):
           return cls(storage_type)
   
   
   storage = pa.UnionArray.from_dense(
       pa.array([0, 1], type=pa.int8()),
       pa.array([0, 0], type=pa.int32()),
       [pa.array([1.5], type=pa.float64()), pa.array(["abc"], 
type=pa.large_string())],
   )
   pa.register_extension_type(MyExt(storage.type))
   ext = pa.ExtensionArray.from_storage(MyExt(storage.type), storage)
   t = pa.table({"x": ext})
   t.validate(full=True)  # passes; t.column("x").to_pylist() == [1.5, "abc"]
   
   sink = io.BytesIO()
   with pa.ipc.new_file(sink, t.schema) as w:
       w.write_table(t)
   t2 = pa.ipc.open_file(io.BytesIO(sink.getvalue())).read_all()
   
   t2.validate(full=True)
   # ArrowInvalid: Column 0: In chunk 0: Invalid: Buffer #1 too small in array 
of
   # type extension<my.ext<MyExt>> and length 2: expected at least 2 byte(s), 
got 0
   
   t2.column("x").to_pylist()
   # Segmentation fault (exit code 139)
   ```
   
   All eight combinations of {dense, sparse} union × {file, stream} IPC format 
× {extension-wrapped, bare}:
   
   | storage | column type | format | result |
   |---|---|---|---|
   | dense union | extension | file | `ArrowInvalid: Buffer #1 too small …` |
   | dense union | extension | stream | `ArrowInvalid: Buffer #1 too small …` |
   | sparse union | extension | file | `ArrowInvalid: Buffer #1 too small …` |
   | sparse union | extension | stream | `ArrowInvalid: Buffer #1 too small …` |
   | dense union | bare | file | OK |
   | dense union | bare | stream | OK |
   | sparse union | bare | file | OK |
   | sparse union | bare | stream | OK |
   
   The defect is on the write side: reading the same file in a fresh process in 
which `my.ext` is NOT registered falls back to the storage type (a plain 
`dense_union` column with `ARROW:extension:name` in the field metadata) and 
fails `validate(full=True)` with the same error, so the union's buffers are 
absent from the serialized bytes themselves rather than being lost during 
extension-type reconstruction.
   
   Reproduced with pyarrow 23.0.1 and 25.0.0 (pip wheels), Python 3.10.20, 
macOS 26.5 arm64. Also observed through `pyarrow.feather` version 2 files, 
which use the IPC file format. The commit history of `cpp/src/arrow/ipc/` 
contains no related fix after 25.0.0.
   
   This looks like the same mechanism as #15068 (IPC roundtrip of an extension 
type with `pa.null()` storage reconstructs invalid buffers): both null and 
union are types whose buffer layout differs from the default — null has no 
buffers and unions have no validity buffer — and in both cases the extension 
wrapper breaks the IPC buffer accounting for them. Possibly one for the #49963 
umbrella.
   
   Context: encountered while property-based-testing the feather I/O of the 
Awkward Array library (scikit-hep/awkward).
   
   ### Component(s)
   
   C++, Python
   
   🤖 Generated with [Claude Code](https://claude.com/claude-code)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to