etseidl opened a new issue, #11207:
URL: https://github.com/apache/arrow-rs/issues/11207

   ### Describe the bug
   
   `ParquetMetaDataReader::load_via_suffix_and_finish` can fail or decode 
incorrect page-index bytes when a prefetch hint returns the footer metadata 
plus preceding file bytes.
   
   `MetadataSuffixFetch` provides only the suffix bytes, not the total file 
size. However, `load_metadata_via_suffix` currently records the preceding bytes 
as beginning at file offset `0`. `load_page_index_with_remainder` then uses 
that incorrect offset when slicing page-index data.
   
   Depending on the requested page-index range, this either returns an error 
such as:
   
   ```text
   Corrupted parquet file: index data range (...) exceeds remainder length (...)
   ```
   
   or may decode bytes belonging to a different part of the file.
   
   ### To Reproduce
   
   Load a file containing page indexes using:
   
   - `PageIndexPolicy::Required`
   - `load_via_suffix_and_finish`
   - A prefetch hint large enough to include the footer metadata and some 
preceding bytes, but smaller than the full file
   
   For example, the existing `alltypes_tiny_pages.parquet` fixture reproduces 
the failure with a prefetch size of `file_len - 1000`.
   
   ### Expected behavior
   
   The page indexes should be fetched using their actual file ranges and 
decoded correctly.
   
   Suffix bytes should only be reused when their absolute file offset is known.
   
   ### Additional context
   
   The suffix-loading API does not expose the total file size, so the absolute 
offset of bytes preceding the footer metadata cannot be determined safely. 
These bytes should not be supplied as a reusable page-index remainder.
   
   This bug was uncovered during a review of #11157 (referenced as `C1`)


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to