etseidl commented on PR #11157:
URL: https://github.com/apache/arrow-rs/pull/11157#issuecomment-5920951864

   > I think this will be more important once people starting to use the 
incremental page index parsing. I think the usecase would go something like:
   > 
   >     1. First query from the file `SELECT ... WHERE a > 5`, reads the 
PageIndex for `a`
   > 
   >     2. Second query from the file `SELECT ... WHERE b > 10`  reads the 
PageIndex for `b`
   > 
   >     3. Second query from the file `SELECT ... WHERE a > 5 AND b > 10`
   > 
   > 
   > Ideally it will be possible / easy to just parse the PageIndex for `b` and 
add it to the page index for `a` so that by the time the third query gets run 
it can reuse the previously parsed index
   
   Ugh, I keep diving into the weeds on this response. TL;DR is I agree, but I 
no longer think the current metadata parser is the place to handle this. The 
parser should parse components, and something else should cache and synthesize 
those parts into a form usable by the rest of the crate to return record 
batches. 
   
   > Another thing to think about is how to (eventually) unify BloomFilter into 
this whole thing -- specifically if the BloomFilter should be part of 
ParquetMetadata or not.
   
   Again, I agree. ParquetMetaData in my mind would eventually wrap a Bloom 
filter provider, which would follow the pattern we settle on for the page 
indexes.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to