etseidl opened a new pull request, #11208: URL: https://github.com/apache/arrow-rs/pull/11208
# Which issue does this PR close? - Closes #11207. # Rationale for this change `load_metadata_via_suffix` treated bytes preceding the footer metadata as beginning at file offset zero, despite `MetadataSuffixFetch` not providing the total file size. Page-index loading could consequently slice the wrong bytes or fail with an out-of-bounds range error. # What changes are included in this PR? - Stop reusing suffix bytes whose absolute file offset is unknown. - Fetch page-index data using its actual file range. - Add a regression test covering page-index loading through `load_via_suffix_and_finish` with a large prefetch hint. # Are these changes tested? Yes. The new regression test fails before the fix with: ```text Corrupted parquet file: index data range (...) exceeds remainder length (...) ``` and passes afterward. The focused async metadata-reader suite passes: ```text cargo test -p parquet --lib 'file::metadata::reader::async_tests' --features arrow,async ``` # Are there any user-facing changes? Page-index loading through `MetadataSuffixFetch` now performs a separate range fetch when the prefetched bytes have no known absolute offset. This prevents incorrect page-index decoding. There are no public API changes. # AI assistance I used OpenAI Codex to help investigate the failure, draft the regression test, and implement the fix. I reviewed and understand the resulting changes and take responsibility for them. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
