zhuqi-lucas commented on PR #11168: URL: https://github.com/apache/arrow-rs/pull/11168#issuecomment-5809545244
cc @etseidl @alamb @adriangb this new design for skip page logic ready for a look. It defers dictionary page decode on the skip path so a column chunk that is skipped end to end never decompresses its dictionary; a chunk that is read after a partial skip pays it on the first data page instead. And some benchmark is pretty good, and also clickbench here show good result: ```rust arrow_reader_clickbench/async/Q21 1.00 82.4±0.68ms ? ?/sec 1.32 109.0±0.71ms ? ?/sec arrow_reader_clickbench/async/Q22 1.00 113.5±2.35ms ? ?/sec 1.25 141.4±0.97ms ? ?/sec ``` End to end through DataFusion (apache/datafusion#25686, clickbench_1 / partitioned / pushdown) also no regression, small improvement for Q25, Q26 and Q39 are consistently 7-11% faster , similar pattern for above arrow clickbench. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
