Rachelint commented on issue #24704:
URL: https://github.com/apache/datafusion/issues/24704#issuecomment-6045248107

   I am thinking about how the block approach might fit with the recent 
aggregation optimizations we’ve been exploring, and wanted to share a couple of 
observations:
   
   - for final aggregation, @jayzhan211 has proposed a promising bucket-based 
approach that could improve performance while offering memory-management 
benefits similar to the block approach.
   https://github.com/apache/datafusion/pull/25724
   
   - for partial aggregation, I’ve seen encouraging results from proactive 
flushing(inspired by part of @jayzhan211 's work in #25724 ). Flushing earlier 
seems to avoid accumulating more state than is beneficial, improving 
performance while keeping the working set small. This could also address much 
of the memory pressure at this stage.
   https://github.com/apache/datafusion/pull/26116
   
   If we pursue both optimizations, I wonder where the block approach would 
offer additional benefits. 


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to