alamb commented on issue #6946: URL: https://github.com/apache/arrow-rs/issues/6946#issuecomment-5836396781
This came up in a discussion @adriangb had today. Basically the idea is to enable reads at the page granularity, while today reads are done at the row group granularity. In my mind, one of the main use cases for the `ScanPlan` API in https://github.com/apache/arrow-rs/pull/10555 is to allow incremental decoding of row groups (aka not having to fetch all the pages for all columns used for the columns in that row group up front) Right now, when the `ParquetPushDecoder` starts to read a row group, it will compute all the pages necessary for either evaluating a pushdown filter OR all the pages for the projection. @adriangb mentioned to me that he thought it was possible to pre-push the data needed into the push decoder and it would not ask for more data until it ran out of them. I want to prove/disprove this by finding or writing a test that demonstrates pushing only the pages needed for the next record batch (not the whole row group) and decoding it if possible Here is some research about how the ranges are computed / needed today <details> Where the ranges for all columns are computed - https://github.com/apache/arrow-rs/blob/2f2360d97a7ec1fb3e35f33c5858fcf9289b4a09/parquet/src/arrow/in_memory_row_group.rs#L59-L134 `InMemoryRowGroup::fetch_ranges` iterates every leaf column and returns either a single range per column chunk or ranges for each page - https://github.com/apache/arrow-rs/blob/2f2360d97a7ec1fb3e35f33c5858fcf9289b4a09/parquet/src/arrow/push_decoder/reader_builder/data.rs#L190-L230 `DataRequestBuilder::build` built from the result of `fetch_ranges` Where the push decoder issues the requests, per row group - https://github.com/apache/arrow-rs/blob/2f2360d97a7ec1fb3e35f33c5858fcf9289b4a09/parquet/src/arrow/push_decoder/reader_builder/mod.rs#L544-L562 the filter path builds a DataRequest using the current predicate's projection. - https://github.com/apache/arrow-rs/blob/2f2360d97a7ec1fb3e35f33c5858fcf9289b4a09/parquet/src/arrow/push_decoder/reader_builder/mod.rs#L720-L731 the output path builds a DataRequest using the final projection. Where decoder refuses to proceed until every range is present - https://github.com/apache/arrow-rs/blob/2f2360d97a7ec1fb3e35f33c5858fcf9289b4a09/parquet/src/arrow/push_decoder/reader_builder/data.rs#L48-L54 DataRequest::needed_ranges returns whichever ranges the buffers do not yet hold. - https://github.com/apache/arrow-rs/blob/2f2360d97a7ec1fb3e35f33c5858fcf9289b4a09/parquet/src/arrow/push_decoder/reader_builder/mod.rs#L582-L594 and https://github.com/apache/arrow-rs/blob/2f2360d97a7ec1fb3e35f33c5858fcf9289b4a09/parquet/src/arrow/push_decoder/reader_builder/mod.rs#L762-L773 return NeedsData while any range is missing. The row group is only materialized once the full set is buffered. </details> -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
