sunchao opened a new issue, #11275:
URL: https://github.com/apache/arrow-rs/issues/11275

   ### Is your feature request related to a problem or challenge?
   
   `ParquetMetaDataPushDecoder` discovers the page-index byte span by walking 
every row group's column metadata. It performs this walk even when both index 
policies are `Skip`, and repeats it each time `try_decode()` returns to 
page-index loading while the requested bytes are still unavailable.
   
   For metadata with many column chunks, these traversals repeat work whose 
result has not changed.
   
   ### Describe the solution you'd like
   
   Return immediately when both index policies are `Skip`. Otherwise, retain 
the discovered range while waiting for its bytes, still checking buffer 
availability on every decode attempt. Invalidate the retained range when either 
index policy changes and consume it before decoding or returning an error.
   
   The requested ranges, decoded metadata, and public APIs should remain 
unchanged. Regression coverage should include repeated polls, clearing buffered 
ranges, and policy changes while a range is pending.
   
   ### Describe alternatives you've considered
   
   Continue recomputing the range on every call. The proposed change retains 
only one optional range per decoder.
   
   ### Additional context
   
   OpenAI Codex assisted with the implementation port, tests, review, and this 
issue description.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to