etseidl opened a new issue, #11207: URL: https://github.com/apache/arrow-rs/issues/11207
### Describe the bug `ParquetMetaDataReader::load_via_suffix_and_finish` can fail or decode incorrect page-index bytes when a prefetch hint returns the footer metadata plus preceding file bytes. `MetadataSuffixFetch` provides only the suffix bytes, not the total file size. However, `load_metadata_via_suffix` currently records the preceding bytes as beginning at file offset `0`. `load_page_index_with_remainder` then uses that incorrect offset when slicing page-index data. Depending on the requested page-index range, this either returns an error such as: ```text Corrupted parquet file: index data range (...) exceeds remainder length (...) ``` or may decode bytes belonging to a different part of the file. ### To Reproduce Load a file containing page indexes using: - `PageIndexPolicy::Required` - `load_via_suffix_and_finish` - A prefetch hint large enough to include the footer metadata and some preceding bytes, but smaller than the full file For example, the existing `alltypes_tiny_pages.parquet` fixture reproduces the failure with a prefetch size of `file_len - 1000`. ### Expected behavior The page indexes should be fetched using their actual file ranges and decoded correctly. Suffix bytes should only be reused when their absolute file offset is known. ### Additional context The suffix-loading API does not expose the total file size, so the absolute offset of bytes preceding the footer metadata cannot be determined safely. These bytes should not be supplied as a reusable page-index remainder. This bug was uncovered during a review of #11157 (referenced as `C1`) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
