letian-jiang commented on PR #3086: URL: https://github.com/apache/drill/pull/3086#issuecomment-5956870771
@cgivre Thanks for pointing me to #3023. I have reviewed it, and I agree that both PRs share the same underlying goal: reusing planning work to reduce the overhead of repeated queries. Before considering #3023 production-ready, I think we would need to establish the conditions for safe plan reuse, even for identical SQL. For example: 1. Time-dependent or non-deterministic expressions, such as date/time and random functions, need special handling when expressions are evaluated or folded during planning. 2. Changes to options that affect planning may invalidate a cached plan. 3. Physical plans contain mutable state, such as node assignments, which should be isolated between executions. 4. Changes to the underlying data files may require rebuilding scan state. 5. Changes to the underlying table schema may invalidate the plan. There is also a practical consideration: production workloads often repeat the same query shape with different literals. Restricting reuse to identical SQL text would substantially limit the benefit for those workloads. My implementation focuses on these correctness requirements through eligibility checks, context and table compatibility checks, and reconstruction of a fresh plan and scan state for each execution. It also preserves typed parameter slots so queries with different literals can reuse compatible cached plans. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
