HippoBaro opened a new issue, #11261:
URL: https://github.com/apache/arrow-rs/issues/11261
### Describe the bug
Writing an Arrow `Dictionary<K, FixedSizeBinary>` with `ArrowWriter` can
produce a Parquet file whose physical schema declares `FIXED_LEN_BYTE_ARRAY`
(FLBA), but whose values use the `BYTE_ARRAY` wire representation.
Parquet dictionary entries use `PLAIN` encoding. For `BYTE_ARRAY`, each
entry has a four-byte little-endian length prefix. For FLBA, entries must
contain only the value bytes: their width is already specified by the schema.
The Arrow writer routes dictionary values of type `FixedSizeBinary` through the
byte-array writer, incorrectly adding those prefixes. The same route also emits
`BYTE_ARRAY`-specific size statistics for an FLBA column.
This breaks interoperability with other Parquet implementations. There is a
corresponding reader issue: reading conforming FLBA files into Arrow
`Dictionary<K, FixedSizeBinary>` uses byte-array decoding rather than the
declared physical type. A Rust Arrow round-trip alone can therefore conceal the
mismatch.
### To Reproduce
```rust
use std::{fs::File, sync::Arc};
use arrow_array::{
Array, DictionaryArray, FixedSizeBinaryArray, Int8Array, RecordBatch,
types::Int8Type,
};
use arrow_schema::{Field, Schema};
use parquet::{
arrow::ArrowWriter,
basic::Compression,
column::page::Page,
file::{
properties::{WriterProperties, WriterVersion},
reader::{FileReader, SerializedFileReader},
},
};
fn main() -> Result<(), Box<dyn std::error::Error>> {
let values = FixedSizeBinaryArray::try_from_iter(
[b"abcd".as_slice(), b"WXYZ".as_slice()].into_iter(),
)?;
let array = DictionaryArray::<Int8Type>::new(
Int8Array::from(vec![0, 1, 0]),
Arc::new(values),
);
let schema = Arc::new(Schema::new(vec![Field::new(
"value", array.data_type().clone(), false,
)]));
let batch = RecordBatch::try_new(schema.clone(), vec![Arc::new(array)])?;
let props = WriterProperties::builder()
.set_writer_version(WriterVersion::PARQUET_2_0)
.set_compression(Compression::UNCOMPRESSED)
.set_dictionary_enabled(true)
.build();
let mut writer =
ArrowWriter::try_new(File::create("rust-flba.parquet")?, schema, Some(props))?;
writer.write(&batch)?;
writer.close()?;
// Inspect the decompressed dictionary independently of Arrow decoding.
let file = SerializedFileReader::new(File::open("rust-flba.parquet")?)?;
let row_group = file.get_row_group(0)?;
let mut pages = row_group.get_column_page_reader(0)?;
if let Some(Page::DictionaryPage { buf, num_values, .. }) =
pages.get_next_page()? {
println!("Dictionary: {num_values} entries, {} bytes", buf.len());
println!("Payload: {:02x?}", buf.as_ref());
} else {
panic!("expected a dictionary page");
}
Ok(())
}
```
Actual output:
```text
Dictionary: 2 entries, 16 bytes
Payload: [04, 00, 00, 00, 61, 62, 63, 64, 04, 00, 00, 00, 57, 58, 59, 5a]
```
The two `04 00 00 00` length prefixes do not belong in an FLBA dictionary.
Two entries of width 4 must occupy exactly **8 bytes**, not 16. This byte-level
check confirms the encoding bug independently of another reader's error
handling.
To also check interoperability using the Arrow C++ reader through PyArrow,
run the following from the same directory (requires Python with a compatible
PyArrow 21.0.0 wheel):
```sh
python3 -m venv /tmp/flba-repro-venv
/tmp/flba-repro-venv/bin/python -m pip install pyarrow==21.0.0
/tmp/flba-repro-venv/bin/python - <<'PY'
import pyarrow.parquet as pq
print(pq.ParquetFile("rust-flba.parquet").schema)
actual = pq.read_table("rust-flba.parquet").column("value").to_pylist()
print(actual)
assert actual == [b"abcd", b"WXYZ", b"abcd"]
PY
```
The schema declares:
```text
required fixed_len_byte_array(4) field_id=-1 value;
```
With the affected writer, PyArrow 21.0.0 fails with:
```text
OSError: Unencoded byte array data bytes does not support
FIXED_LEN_BYTE_ARRAY
```
### Expected behavior
- The writer should honor the declared `FIXED_LEN_BYTE_ARRAY` physical type,
including dictionary entries and fallback data pages, and should not emit
byte-array-only size statistics for FLBA columns.
- For the example above, the dictionary should contain exactly these 8
bytes, and the PyArrow assertion should pass:
```text
Dictionary: 2 entries, 8 bytes
Payload: [61, 62, 63, 64, 57, 58, 59, 5a]
```
- Conforming FLBA files from other implementations should also decode
correctly when an Arrow dictionary representation is requested.
- Readers should reject malformed FLBA payloads rather than silently
interpreting length prefixes as values or accepting trailing bytes. Dictionary
payloads must contain exactly `num_values * type_length` bytes; PLAIN data-page
value sections must match the physical non-null value count, excluding
repetition and definition levels. Partial reads and skips should not bypass
these checks.
### Additional context
NA
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]