etseidl opened a new pull request, #11208:
URL: https://github.com/apache/arrow-rs/pull/11208

   # Which issue does this PR close?
   
   - Closes #11207.
   
   # Rationale for this change
   
   `load_metadata_via_suffix` treated bytes preceding the footer metadata as 
beginning at file offset zero, despite `MetadataSuffixFetch` not providing the 
total file size. Page-index loading could consequently slice the wrong bytes or 
fail with an out-of-bounds range error.
   
   # What changes are included in this PR?
   
   - Stop reusing suffix bytes whose absolute file offset is unknown.
   - Fetch page-index data using its actual file range.
   - Add a regression test covering page-index loading through 
`load_via_suffix_and_finish` with a large prefetch hint.
   
   # Are these changes tested?
   
   Yes. The new regression test fails before the fix with:
   
   ```text
   Corrupted parquet file: index data range (...) exceeds remainder length (...)
   ```
   
   and passes afterward.
   
   The focused async metadata-reader suite passes:
   
   ```text
   cargo test -p parquet --lib 'file::metadata::reader::async_tests' --features 
arrow,async
   ```
   
   # Are there any user-facing changes?
   
   Page-index loading through `MetadataSuffixFetch` now performs a separate 
range fetch when the prefetched bytes have no known absolute offset. This 
prevents incorrect page-index decoding. There are no public API changes.
   
   # AI assistance
   
   I used OpenAI Codex to help investigate the failure, draft the regression 
test, and implement the fix. I reviewed and understand the resulting changes 
and take responsibility for them.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to