etseidl commented on PR #11157: URL: https://github.com/apache/arrow-rs/pull/11157#issuecomment-5920951864
> I think this will be more important once people starting to use the incremental page index parsing. I think the usecase would go something like: > > 1. First query from the file `SELECT ... WHERE a > 5`, reads the PageIndex for `a` > > 2. Second query from the file `SELECT ... WHERE b > 10` reads the PageIndex for `b` > > 3. Second query from the file `SELECT ... WHERE a > 5 AND b > 10` > > > Ideally it will be possible / easy to just parse the PageIndex for `b` and add it to the page index for `a` so that by the time the third query gets run it can reuse the previously parsed index Ugh, I keep diving into the weeds on this response. TL;DR is I agree, but I no longer think the current metadata parser is the place to handle this. The parser should parse components, and something else should cache and synthesize those parts into a form usable by the rest of the crate to return record batches. > Another thing to think about is how to (eventually) unify BloomFilter into this whole thing -- specifically if the BloomFilter should be part of ParquetMetadata or not. Again, I agree. ParquetMetaData in my mind would eventually wrap a Bloom filter provider, which would follow the pattern we settle on for the page indexes. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
