alamb commented on issue #6946:
URL: https://github.com/apache/arrow-rs/issues/6946#issuecomment-5836396781

   This came up in a discussion @adriangb had today. Basically the idea is to 
enable reads at the page granularity, while today reads are done at the row 
group granularity. 
   
   In my mind, one of the main use cases for the `ScanPlan` API in  
https://github.com/apache/arrow-rs/pull/10555  is to allow incremental decoding 
of row groups (aka not having to fetch all the pages for all columns used for 
the columns in that row group up front)
   
   Right now, when the `ParquetPushDecoder` starts to read a row group, it will 
compute all the pages necessary for either evaluating a pushdown filter OR all 
the pages for the projection. @adriangb  mentioned to me that he thought it was 
possible to pre-push the data needed into the push decoder and it would not ask 
for more data until it ran out of them.
   
   I want to prove/disprove this by finding or writing a test that demonstrates 
pushing only the pages needed for  the next record batch (not the whole row 
group) and decoding it if possible
   
   Here is some research about how the ranges are computed / needed today
   <details>
   
   Where the ranges for all columns are computed
   
   - 
https://github.com/apache/arrow-rs/blob/2f2360d97a7ec1fb3e35f33c5858fcf9289b4a09/parquet/src/arrow/in_memory_row_group.rs#L59-L134
 `InMemoryRowGroup::fetch_ranges` iterates every leaf column and returns either 
a single range per column chunk or ranges for each page
   - 
https://github.com/apache/arrow-rs/blob/2f2360d97a7ec1fb3e35f33c5858fcf9289b4a09/parquet/src/arrow/push_decoder/reader_builder/data.rs#L190-L230
 `DataRequestBuilder::build` built from the result of `fetch_ranges`
   
   Where the push decoder issues the requests, per row group
   
   - 
https://github.com/apache/arrow-rs/blob/2f2360d97a7ec1fb3e35f33c5858fcf9289b4a09/parquet/src/arrow/push_decoder/reader_builder/mod.rs#L544-L562
 the filter path builds a DataRequest using the current predicate's projection.
   - 
https://github.com/apache/arrow-rs/blob/2f2360d97a7ec1fb3e35f33c5858fcf9289b4a09/parquet/src/arrow/push_decoder/reader_builder/mod.rs#L720-L731
 the output path builds a DataRequest using the final projection.
   
   Where decoder refuses to proceed until every range is present
   
   - 
https://github.com/apache/arrow-rs/blob/2f2360d97a7ec1fb3e35f33c5858fcf9289b4a09/parquet/src/arrow/push_decoder/reader_builder/data.rs#L48-L54
 DataRequest::needed_ranges returns whichever ranges the buffers do not yet 
hold.
   - 
https://github.com/apache/arrow-rs/blob/2f2360d97a7ec1fb3e35f33c5858fcf9289b4a09/parquet/src/arrow/push_decoder/reader_builder/mod.rs#L582-L594
 and 
https://github.com/apache/arrow-rs/blob/2f2360d97a7ec1fb3e35f33c5858fcf9289b4a09/parquet/src/arrow/push_decoder/reader_builder/mod.rs#L762-L773
 return NeedsData while any range is missing. The row group is only 
materialized once the full set is buffered.
   
   </details>


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to