HippoBaro opened a new pull request, #11262: URL: https://github.com/apache/arrow-rs/pull/11262
# Which issue does this PR close? <!-- We generally require a GitHub issue to be filed for all bug fixes and enhancements and this helps us generate change logs for our releases. You can link an issue to this PR using the GitHub syntax. --> - Closes #11261 # Rationale for this change <!-- Why are you proposing this change? If this is already explained clearly in the issue then this section is not needed. Explaining clearly why changes are proposed helps reviewers understand your changes and offer better suggestions for fixes. --> The writer can encode `Dictionary<K, FixedSizeBinary>` values using the `BYTE_ARRAY` wire representation even though the Parquet schema declares `FIXED_LEN_BYTE_ARRAY` (FLBA). This distinction matters for PLAIN encoding, which is also used for dictionary entries: `BYTE_ARRAY` prefixes each value with a four-byte length, whereas FLBA stores only the value bytes because their width is defined by the schema. # What changes are included in this PR? <!-- There is no need to duplicate the description in the issue here but it is sometimes worth providing a summary of the individual changes in this PR. --> - Route Arrow `Dictionary<K, FixedSizeBinary>` values through the FLBA writer rather than the byte-array writer. - Align dictionary decoding, fallback decoding, and Arrow dictionary spill conversion with the declared physical type. - Stop emitting `BYTE_ARRAY`-specific size statistics for FLBA columns. - Require decompressed FLBA dictionary payloads to contain exactly `num_values * type_length` bytes, using checked arithmetic, in the typed column reader and both Arrow dictionary-decoding paths. - Validate PLAIN data-page value sections independently of repetition and definition levels, counting physical non-null values. - Validate affected pages before exposing their values, including partial reads, skips, and PLAIN pages following dictionary-encoded pages. The reader follows the schema rather than attempting to infer which historical writer produced a payload. # Are these changes tested? Yes. Regression coverage exercises FLBA dictionary handling and malformed dictionary/data-page payloads, including partial reads, skips, and dictionary-to-PLAIN transitions. # Are there any user-facing changes? Yes. This is an **intentional file-compatibility break for malformed historical payloads** produced by the affected dictionary-writing path. Those payloads are no longer accepted as valid FLBA data, and reading them result is a well-defined error. New writes use the physical representation declared by the schema. Already-conforming files remain supported, including files produced by the dense `FixedSizeBinary` writer path. There are no public Rust API signature changes. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
