Rachelint commented on PR #25567: URL: https://github.com/apache/datafusion/pull/25567#issuecomment-5883601349
# String group keys queries It seems q33 and q34 in benchmark machine in #25724 . Seems it is harder to measure when bucket approach will get improvement in string cases(I also found q33 and q34 get improvement in my local, but no change in benchmark machine and my cloud machine...), and it always depends on the hardware... # Data buffer reusing for stringview I agree with `the coalescer's garbage-collection copy is not waste`. Actually I have made simple experiment about directly reusing the data buffer in q33 and q34 in #23168 before, and found no obvious improvement in the benchmark machine(but seems in apple m4, q34 have improvement for buffer reusing?). The root cause may be complex, but it seems we should just copy rather than reusing in most cases... # Q11/Q12/Q14 I think we can check if we encounter `Poll::Pending` in `final aggr`, for understanding if the regression is caused by pipeline blocking? It seems the same problem as q33 and q34, it is hard to measure will bucket approach can get improvement in such cases, and hard to decide when should we enable the bucket approach. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
