HippoBaro opened a new issue, #11261:
URL: https://github.com/apache/arrow-rs/issues/11261

   ### Describe the bug
   
   Writing an Arrow `Dictionary<K, FixedSizeBinary>` with `ArrowWriter` can 
produce a Parquet file whose physical schema declares `FIXED_LEN_BYTE_ARRAY` 
(FLBA), but whose values use the `BYTE_ARRAY` wire representation.
   
   Parquet dictionary entries use `PLAIN` encoding. For `BYTE_ARRAY`, each 
entry has a four-byte little-endian length prefix. For FLBA, entries must 
contain only the value bytes: their width is already specified by the schema. 
The Arrow writer routes dictionary values of type `FixedSizeBinary` through the 
byte-array writer, incorrectly adding those prefixes. The same route also emits 
`BYTE_ARRAY`-specific size statistics for an FLBA column.
   
   This breaks interoperability with other Parquet implementations. There is a 
corresponding reader issue: reading conforming FLBA files into Arrow 
`Dictionary<K, FixedSizeBinary>` uses byte-array decoding rather than the 
declared physical type. A Rust Arrow round-trip alone can therefore conceal the 
mismatch.
   
   
   ### To Reproduce
   
   ```rust
   use std::{fs::File, sync::Arc};
   
   use arrow_array::{
       Array, DictionaryArray, FixedSizeBinaryArray, Int8Array, RecordBatch, 
types::Int8Type,
   };
   use arrow_schema::{Field, Schema};
   use parquet::{
       arrow::ArrowWriter,
       basic::Compression,
       column::page::Page,
       file::{
           properties::{WriterProperties, WriterVersion},
           reader::{FileReader, SerializedFileReader},
       },
   };
   
   fn main() -> Result<(), Box<dyn std::error::Error>> {
       let values = FixedSizeBinaryArray::try_from_iter(
           [b"abcd".as_slice(), b"WXYZ".as_slice()].into_iter(),
       )?;
       let array = DictionaryArray::<Int8Type>::new(
           Int8Array::from(vec![0, 1, 0]),
           Arc::new(values),
       );
       let schema = Arc::new(Schema::new(vec![Field::new(
           "value", array.data_type().clone(), false,
       )]));
       let batch = RecordBatch::try_new(schema.clone(), vec![Arc::new(array)])?;
       let props = WriterProperties::builder()
           .set_writer_version(WriterVersion::PARQUET_2_0)
           .set_compression(Compression::UNCOMPRESSED)
           .set_dictionary_enabled(true)
           .build();
       let mut writer = 
ArrowWriter::try_new(File::create("rust-flba.parquet")?, schema, Some(props))?;
       writer.write(&batch)?;
       writer.close()?;
   
       // Inspect the decompressed dictionary independently of Arrow decoding.
       let file = SerializedFileReader::new(File::open("rust-flba.parquet")?)?;
       let row_group = file.get_row_group(0)?;
       let mut pages = row_group.get_column_page_reader(0)?;
       if let Some(Page::DictionaryPage { buf, num_values, .. }) = 
pages.get_next_page()? {
           println!("Dictionary: {num_values} entries, {} bytes", buf.len());
           println!("Payload: {:02x?}", buf.as_ref());
       } else {
           panic!("expected a dictionary page");
       }
       Ok(())
   }
   ```
   
   Actual output:
   
   ```text
   Dictionary: 2 entries, 16 bytes
   Payload: [04, 00, 00, 00, 61, 62, 63, 64, 04, 00, 00, 00, 57, 58, 59, 5a]
   ```
   
   The two `04 00 00 00` length prefixes do not belong in an FLBA dictionary. 
Two entries of width 4 must occupy exactly **8 bytes**, not 16. This byte-level 
check confirms the encoding bug independently of another reader's error 
handling.
   
   To also check interoperability using the Arrow C++ reader through PyArrow, 
run the following from the same directory (requires Python with a compatible 
PyArrow 21.0.0 wheel):
   
   ```sh
   python3 -m venv /tmp/flba-repro-venv
   /tmp/flba-repro-venv/bin/python -m pip install pyarrow==21.0.0
   /tmp/flba-repro-venv/bin/python - <<'PY'
   import pyarrow.parquet as pq
   
   print(pq.ParquetFile("rust-flba.parquet").schema)
   actual = pq.read_table("rust-flba.parquet").column("value").to_pylist()
   print(actual)
   assert actual == [b"abcd", b"WXYZ", b"abcd"]
   PY
   ```
   
   The schema declares:
   
   ```text
   required fixed_len_byte_array(4) field_id=-1 value;
   ```
   
   With the affected writer, PyArrow 21.0.0 fails with:
   
   ```text
   OSError: Unencoded byte array data bytes does not support 
FIXED_LEN_BYTE_ARRAY
   ```
   
   
   ### Expected behavior
   
   - The writer should honor the declared `FIXED_LEN_BYTE_ARRAY` physical type, 
including dictionary entries and fallback data pages, and should not emit 
byte-array-only size statistics for FLBA columns.
   - For the example above, the dictionary should contain exactly these 8 
bytes, and the PyArrow assertion should pass:
   
     ```text
     Dictionary: 2 entries, 8 bytes
     Payload: [61, 62, 63, 64, 57, 58, 59, 5a]
     ```
   
   - Conforming FLBA files from other implementations should also decode 
correctly when an Arrow dictionary representation is requested.
   - Readers should reject malformed FLBA payloads rather than silently 
interpreting length prefixes as values or accepting trailing bytes. Dictionary 
payloads must contain exactly `num_values * type_length` bytes; PLAIN data-page 
value sections must match the physical non-null value count, excluding 
repetition and definition levels. Partial reads and skips should not bypass 
these checks.
   
   
   ### Additional context
   
   NA


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to