alamb opened a new pull request, #11218:
URL: https://github.com/apache/arrow-rs/pull/11218

   # Which issue does this PR close?
   
   - Related to https://github.com/apache/arrow-rs/issues/6946
   
   # Rationale for this change
   
   While discussing https://github.com/apache/arrow-rs/issues/6946 (reading at 
page rather than row group granularity), a question came up about whether 
`ParquetPushDecoder` would decode a batch if only the pages needed for that 
batch had been pushed, and only ask for more data once it ran out.
   
   This PR adds a test that answers that question for the current code: it does 
not.
   
   # What changes are included in this PR?
   
   A new test, `test_decoder_first_page_only_does_not_decode`, plus a helper 
that loads the offset index into the test file metadata. The test uses the 
existing test file (2 row groups of 200 rows, 100 rows per data page) with a 
batch size of 100, so the first batch of a row group needs only the first data 
page of each column.
   
   The test shows:
   
   1. When the decoder starts row group 0 it requests the entire column chunk 
for every projected column (`DataRequestBuilder::build` via 
`InMemoryRowGroup::fetch_ranges`).
   2. After pushing only the dictionary page and first data page of each column 
(enough to decode the first batch), `try_decode` still returns `NeedsData` for 
the original full ranges and produces no batch.
   3. After additionally pushing the second data page of each column as 
separate ranges, `try_decode` *still* returns `NeedsData` for the original 
ranges. `PushBuffers::has_range` only considers a range satisfied when a single 
pushed buffer covers it entirely; adjacent pushed ranges are not coalesced.
   4. Only once the exact ranges originally requested are pushed does the 
decoder produce batches.
   
   No non-test code is changed.
   
   # Are these changes tested?
   
   This PR is only a test.
   
   # Are there any user-facing changes?
   
   No.
   
   🤖 Generated with [Claude Code](https://claude.com/claude-code)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to