Hello Riza Suminto, Csaba Ringhofer, Impala Public Jenkins,
I'd like you to reexamine a change. Please visit
http://gerrit.cloudera.org:8080/24883
to look at the new patch set (#2).
Change subject: IMPALA-13052: Estimate reservoir sample memory in aggregations
......................................................................
IMPALA-13052: Estimate reservoir sample memory in aggregations
sample(), appx_median() and histogram() keep a ReservoirSampleState per
group, allocated by FunctionContext::Allocate() outside the tuple and
the reservation, but the planner counted only the 12-byte StringValue
slot. The state with its initial sample array takes about 600 bytes per
call and group even for one row, and up to 1.3MB (2.3MB for DECIMAL) at
20000 samples. For the grouping query in the Jira the merge aggregation
was estimated at 18MB; it is now estimated at 312MB per instance, and
three instances used 0.90GB together.
AggregationNode now adds this memory, mirroring the state and its
FreePool and MemPool allocations; static asserts in
aggregate-functions-ir.cc fail the build if the state or sample sizes
change. The FreePool part is not specific to aggregations and goes to
PlannerContext.getFreePoolBytes(). The estimate takes the rows per group
from the first phase so that merges are covered. Running short of this
memory never makes an aggregation spill (IMPALA-3304), so it is added
after the caps that assume spilling. On lineitem the estimate was within
1% of the measured memory at 4, 600 and 60K rows per group before
IMPALA-15381, and at 4 rows it still is: 592 bytes against 598.
A streaming preaggregation keeps the planner's usual group count,
bounded by PREAGG_BYTES_LIMIT when set, which errs high: 1.5M groups per
instance for the Jira query, where each instance holds about 500K. The
backend's limit on hash table growth (STREAMING_HT_MIN_REDUCTION) is not
mirrored, because it depends on how the input of one instance reduces.
The planner expects a reduction of 1.33 both for the Jira query, which
reduced fourfold and kept all its groups, and for a COUNT(DISTINCT) over
unique rows, whose preaggregation stopped at 98,304 groups.
Execution and the row size of the intermediate tuple are unchanged;
admission control sees larger estimates for queries with these
functions.
Testing:
- New reservoir-sample-agg.test: the Jira examples, groups at 20000
samples, COUNT(DISTINCT), MEM_ESTIMATE_SCALE_FOR_SPILLING_OPERATOR and
PREAGG_BYTES_LIMIT. Reverting each part of the estimate separately
fails the cases that part covers.
- TpcdsPlannerTest and nine aggregation planner test files produce the
same plans. With PLANNER=CALCITE the new estimates are the same.
Change-Id: Id6eec3daf7fea9b027ff0be9c95c6ee5a3cf323d
Assisted-by: claude-opus-5 (Claude Code)
---
M be/src/exprs/aggregate-functions-ir.cc
M fe/src/main/java/org/apache/impala/analysis/FunctionCallExpr.java
M fe/src/main/java/org/apache/impala/planner/AggregationNode.java
M fe/src/main/java/org/apache/impala/planner/PlannerContext.java
M fe/src/test/java/org/apache/impala/planner/PlannerTest.java
A
testdata/workloads/functional-planner/queries/PlannerTest/reservoir-sample-agg.test
6 files changed, 618 insertions(+), 5 deletions(-)
git pull ssh://gerrit.cloudera.org:29418/Impala-ASF refs/changes/83/24883/2
--
To view, visit http://gerrit.cloudera.org:8080/24883
To unsubscribe, visit http://gerrit.cloudera.org:8080/settings
Gerrit-Project: Impala-ASF
Gerrit-Branch: master
Gerrit-MessageType: newpatchset
Gerrit-Change-Id: Id6eec3daf7fea9b027ff0be9c95c6ee5a3cf323d
Gerrit-Change-Number: 24883
Gerrit-PatchSet: 2
Gerrit-Owner: Aleksandr Efimov <[email protected]>
Gerrit-Reviewer: Aleksandr Efimov <[email protected]>
Gerrit-Reviewer: Csaba Ringhofer <[email protected]>
Gerrit-Reviewer: Impala Public Jenkins <[email protected]>
Gerrit-Reviewer: Riza Suminto <[email protected]>