alamb commented on PR #11157: URL: https://github.com/apache/arrow-rs/pull/11157#issuecomment-5919939770
> As C11 points out, we need to do something and document it. The previous behavior was not a contract, just a consequence of how things were implemented. 60.0.0 introduced a behavior change that we should either revert or document anyway. I guess at this point I'm still open to either path, either what I have now (always replace), or go back to preserve across multiple calls, and take that into account as we try to make the back-end storage more efficient. Sounds like @alamb is a vote for the latter; I abstain. Other votes? @adriangb, @zhuqi-lucas, @sunchao? Maybe we are worrying about old behavior that no one really cares about 🤔 I think this will be more important once people starting to use the incremental page index parsing. I think the usecase would go something like: 1. First query from the file `SELECT ... WHERE a > 5`, reads the PageIndex for `a` 2. Second query from the file `SELECT ... WHERE b > 10` reads the PageIndex for `b` 2. Second query from the file `SELECT ... WHERE a > 5 AND b > 10` Ideally it will be possible / easy to just parse the PageIndex for `b` and add it to the page index for `a` so that by the time the third query gets run it can reuse the previously parsed index -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
