adriangb opened a new issue, #11234:
URL: https://github.com/apache/arrow-rs/issues/11234

   Part of https://github.com/apache/arrow-rs/issues/6946.
   
   # Goal
   
   Let `ParquetPushDecoder` request and decode one batch at a time, so that a 
caller gets the first batch of a row group as soon as the pages of that batch 
are pushed, and the decoder holds approximately one batch of bytes, not the 
full row group.
   
   ```rust
   let mut decoder = ParquetPushDecoderBuilder::try_new_decoder(metadata)?
       .with_fetch_granularity(FetchGranularity::Batch) // new, opt-in
       .build()?;
   // NeedsData asks only for the pages of the next batch.
   ```
   
   Prototype measurement (one row group of 84 MB, 50 ms request latency):
   
   | Mode | First batch | Peak buffered bytes |
   |---|---:|---:|
   | `FetchGranularity::RowGroup` (default) | 451 ms | 84 MB |
   | `FetchGranularity::Batch`, read-ahead of 8 MB | 89 ms | 8.1 MB |
   
   https://github.com/apache/arrow-rs/pull/11223 has the full change (+4,156 
lines). It is too large to review in one step, so this EPIC splits it.
   
   # Plan
   
   Each PR is useful without the PRs after it. No PR adds a temporary error or 
fallback that a later PR removes. A later PR can add an optimization to code 
that is already correct.
   
   | PR | Content | Base | Useful alone because |
   |---|---|---|---|
   | **A** https://github.com/apache/arrow-rs/pull/11218 | Test: the decoder 
needs the full row group before it returns a batch | `main` | Records the 
current behavior |
   | **C** TBD | `into_builder`: release the pushed bytes of row groups that 
the new decoder does not read | `main` | Today these bytes stay until 
`clear_all_ranges` |
   | **D** TBD | Keep `PushBuffers` sorted by offset | `main` | Lookups become 
a binary search. Today each lookup scans all buffers. |
   | **E** TBD | `FetchGranularity::Batch`, with and without a `RowFilter`. 
Holds the bytes of a row group until the row group ends. | A, #11233 fix | 
Lower time to the first batch |
   | **F** TBD | Batch mode: release each data page when all readers have 
passed it | E, D | Lower peak memory |
   | **G** TBD | Batch mode: use the predicate cache | E | Filtered scans do 
not decode predicate columns two times |
   | **H** TBD | Batch mode: use the configured `RowSelectionPolicy` (mask) in 
each window | E | Faster dense selections |
   
   ```text
   main ──┬── A  test ─────────────────────┐
          ├── #11233 fix ──────────────────┴── E  batch granularity ──┬── F  
page release ◀── also needs D
          ├── C  into_builder release                                 ├── G  
predicate cache
          └── D  sorted PushBuffers ──────────────────────────────────┴── H  
selection policy
   ```
   
   A, C, D and the #11233 fix are independent and can be reviewed in parallel. 
F, G and H are independent of each other.
   
   # Tangential
   
   - https://github.com/apache/arrow-rs/issues/11233: the decoder does not 
release a pushed buffer that is larger than the requested ranges. This is a bug 
in the default mode. E uses the fix (`PushBuffers::release_ranges`), but the 
fix is not part of this EPIC.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to