Aleksandr Efimov has uploaded this change for review. ( 
http://gerrit.cloudera.org:8080/24883


Change subject: IMPALA-13052: Estimate reservoir sample memory in aggregations
......................................................................

IMPALA-13052: Estimate reservoir sample memory in aggregations

sample(), appx_median() and histogram() keep a ReservoirSampleState per
group, allocated by FunctionContext::Allocate() outside the tuple and
the reservation, but the planner counted only the 12-byte StringValue
slot. The state takes about 4.5KB per call and group even for one row,
most of it the mt19937_64 generator, and up to 1.3MB (2.3MB for DECIMAL)
at 20000 samples. For the grouping query in the Jira the merge
aggregation was estimated at 18MB and used 2.19GB; it is now estimated
at 2.27GB.

AggregationNode now adds this memory, mirroring the state and its
FreePool and MemPool allocations; static asserts in
aggregate-functions-ir.cc fail the build if the state or sample sizes
change. It takes the rows per group from the first phase so that merges
are covered. Running short of this memory never makes an aggregation
spill (IMPALA-3304), so it is added after the caps that assume spilling.
On lineitem the estimate is within 1% of the measured memory at 4, 600
and 60K rows per group.

A streaming preaggregation keeps the planner's usual group count,
bounded by PREAGG_BYTES_LIMIT when set, which errs high: 1.5M groups per
instance for the Jira query, where each instance holds about 500K. The
backend's limit on hash table growth (STREAMING_HT_MIN_REDUCTION) is not
mirrored, because it depends on how the input of one instance reduces.
The planner expects a reduction of 1.33 both for the Jira query, which
reduced fourfold and kept all its groups, and for a COUNT(DISTINCT) over
unique rows, whose preaggregation stopped at 98,304 groups.

Execution and the row size of the intermediate tuple are unchanged;
admission control sees larger estimates for queries with these
functions.

Testing:
- New reservoir-sample-agg.test: the Jira examples, groups at 20000
  samples, COUNT(DISTINCT), MEM_ESTIMATE_SCALE_FOR_SPILLING_OPERATOR and
  PREAGG_BYTES_LIMIT. Reverting each part of the estimate separately
  fails the cases that part covers.
- TpcdsPlannerTest and nine aggregation planner test files produce the
  same plans. With PLANNER=CALCITE the new estimates are the same.

Change-Id: Id6eec3daf7fea9b027ff0be9c95c6ee5a3cf323d
Assisted-by: claude-opus-5 (Claude Code)
---
M be/src/exprs/aggregate-functions-ir.cc
M fe/src/main/java/org/apache/impala/analysis/FunctionCallExpr.java
M fe/src/main/java/org/apache/impala/planner/AggregationNode.java
M fe/src/test/java/org/apache/impala/planner/PlannerTest.java
A 
testdata/workloads/functional-planner/queries/PlannerTest/reservoir-sample-agg.test
5 files changed, 618 insertions(+), 5 deletions(-)



  git pull ssh://gerrit.cloudera.org:29418/Impala-ASF refs/changes/83/24883/1
--
To view, visit http://gerrit.cloudera.org:8080/24883
To unsubscribe, visit http://gerrit.cloudera.org:8080/settings

Gerrit-Project: Impala-ASF
Gerrit-Branch: master
Gerrit-MessageType: newchange
Gerrit-Change-Id: Id6eec3daf7fea9b027ff0be9c95c6ee5a3cf323d
Gerrit-Change-Number: 24883
Gerrit-PatchSet: 1
Gerrit-Owner: Aleksandr Efimov <[email protected]>

Reply via email to