HippoBaro opened a new pull request, #11278:
URL: https://github.com/apache/arrow-rs/pull/11278

   # Which issue does this PR close?
   
   - Closes #11277.
   
   # Rationale for this change
   
   The writer produces schemas that do not always match the logical data it 
stores.
   
   The main case involves run-end-encoded arrays. REE arrays have no parent 
validity bitmap and report a physical null count of zero even when their run 
values contain nulls. The writer can therefore declare a non-nullable outer REE 
field as a required Parquet column, then encode nullable run values without the 
definition levels needed to represent those nulls. This silently replaces nulls 
with ordinary values.
   
   Embedded Arrow metadata has a related consistency problem. It can preserve 
dictionary wrappers around nested values that the Rust reader cannot 
reconstruct. The physical Parquet values and levels are valid in these cases, 
but `ARROW:schema` describes a representation that prevents the default reader 
from opening the file.
   
   # What changes are included in this PR?
   
   - Recursively resolve REE value types when deriving the Parquet schema and 
remove REE wrappers from embedded Arrow metadata, while preserving field names 
and metadata.
   - Merge enclosing-field and run-value nullability so nullable run values use 
optional Parquet fields and the corresponding definition levels.
   - Preserve supported scalar dictionary hints, including dictionaries of 
`Utf8View` and `BinaryView`, while removing dictionary wrappers around nested 
values that the reader cannot reconstruct.
   - Recurse through dictionary value types when constructing column writers.
   - Generate nested levels from the writer's target child fields rather than 
from child declarations on the incoming array.
   - Compare compatible inputs by logical value type, keeping the writer's 
declared schema contract independent of alternate dense, dictionary, run-end, 
offset-width, and view layouts.
   - Validate resolved map-key nullability during Arrow-to-Parquet schema 
conversion, rejecting key schemas that are nullable through their encoding 
wrappers.
   
   # Are these changes tested?
   
   Yes. Regression coverage includes:
   
   - Recursive REE schema conversion and propagation of run-value nullability.
   - Nested metadata normalization and preservation of supported scalar 
view-dictionary hints.
   - Logically nullable map-key schema rejection, with controls for required 
keys and nullable members inside valid key structures.
   - Dense batches written under compatible nested dictionary/REE schemas, 
including struct, list, and map children.
   
   # Are there any user-facing changes?
   
   Yes. This is a **correctness-driven schema compatibility change**:
   
   - Affected REE fields may become optional where they were previously 
declared required, allowing their logical nulls to be represented correctly.
   - Embedded Arrow metadata is normalized to reader-supported representations. 
Unsupported nested dictionary wrappers are no longer preserved, while supported 
scalar dictionary hints remain intact.
   - Arrow-to-Parquet schema conversion rejects map-key schemas whose resolved 
keys are nullable instead of emitting nullable Parquet map keys.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to