yashmayya opened a new pull request, #19745:
URL: https://github.com/apache/pinot/pull/19745

   ## Summary
   
   Planning queries with large IN lists has regressed more than once without 
any test failing. The last time, a Calcite upgrade and new planner rules made a 
common query shape about 3.7x slower to plan, and the regression was found in 
production. #19674 fixed it. This PR adds tests that fail if these shapes get 
expensive again, for example after a Calcite upgrade or a new planner rule.
   
   ## How the tests measure cost
   
   `LargeInListPlanningBudgetTest` plans each shape and compares its cost with 
a baseline query that holds the same IN lists in a single-table `WHERE`:
   
   - **Planner work:** the bytes that the planning thread allocates 
(`ThreadMXBean#getCurrentThreadAllocatedBytes`). The absolute bytes change with 
the JIT state, but the ratio to the baseline moved by at most 0.05 across C2, 
C1 only, no escape analysis, and the JaCoCo agent. So the budgets are ratios. 
The tests do not use time, because time depends on the machine and its load.
   - **Plan size:** the serialized bytes of the stages.
   
   | Shapes | Values | Budget | With #19674 | Before #19674 |
   |---|---|---|---|---|
   | 4 join shapes: 2 or 3 LEFT or INNER JOINs, 3 IN or NOT IN lists, GROUP BY, 
FILTER aggregates, ORDER BY LIMIT | 10,000 per list | 1.5x | 1.0x–1.1x | 
2.8x–3.6x |
   | 9 lists outside `WHERE`: JOIN ON, LEFT JOIN ON, CASE, CASE with NOT IN, 
FILTER, SELECT list, GROUP BY, window, CASE in HAVING | 300 | 5x | 0.9x–1.7x | 
387x–404x |
   | 2 lists next to a range, plan bytes: IN or a range, NOT IN and a range | 
20,000 | 2x | 1.0x | 9.6x–11.9x |
   
   Two more tests keep the budgets meaningful:
   
   - `testBudgetsFailWithoutSealing` plans 2 shapes with 
`sealedInListThreshold=0` and checks that they exceed the budgets. It fails if 
the budgets stop catching the regression.
   - `testBaselineGrowsLinearly` checks that the baseline allocates in 
proportion to the number of values, because a ratio cannot see a cost that the 
baseline has too.
   
   ## Verification
   
   - Current master (with #19674): 17 of 17 pass, in about 13 s.
   - Master before #19674 (ff069f8ca8): all 15 budget tests fail.
   - #19674 with `DEFAULT_SEALED_IN_LIST_THRESHOLD = 0`: the 13 allocation 
tests fail. The 2 plan-size tests pass, because the IN/range split in 
`RexExpressionUtils` does not depend on the threshold.
   
   ## Not covered
   
   - Many small IN lists, or a long `x = 1 OR x = 2 ...` chain, inside CASE or 
JOIN ON. These still plan in cubic time (a follow-up of #19674), so they have 
no budget yet.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to