Rachelint commented on issue #24704: URL: https://github.com/apache/datafusion/issues/24704#issuecomment-6045248107
I am thinking about how the block approach might fit with the recent aggregation optimizations we’ve been exploring, and wanted to share a couple of observations: - for final aggregation, @jayzhan211 has proposed a promising bucket-based approach that could improve performance while offering memory-management benefits similar to the block approach. https://github.com/apache/datafusion/pull/25724 - for partial aggregation, I’ve seen encouraging results from proactive flushing(inspired by part of @jayzhan211 's work in #25724 ). Flushing earlier seems to avoid accumulating more state than is beneficial, improving performance while keeping the working set small. This could also address much of the memory pressure at this stage. https://github.com/apache/datafusion/pull/26116 If we pursue both optimizations, I wonder where the block approach would offer additional benefits. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
