[email protected] has uploaded this change for review. (
http://gerrit.cloudera.org:8080/24925
Change subject: IMPALA-9976: Recover admission state after an admissiond restart
......................................................................
IMPALA-9976: Recover admission state after an admissiond restart
When the admissiond restarts it loses the state of all queries it
admitted or queued. Running queries keep running, but the new admissiond
does not count them against pool and host limits, so it over-admits
until they finish, and their later ReleaseQuery rpcs fail. Queued
queries fail once their coordinator's GetQueryStatus rpc can no longer
reach the admissiond or the new one does not know the query.
The statestore cannot be used to catch up, as the JIRA suggested: the
request-queue topic only carries per-host and per-pool aggregates, which
it deletes when the admissiond that published them fails, and the
per-query allocations needed to release a query correctly are only
known to the admissiond and to the query's coordinator. This change
makes the coordinators the source of truth:
- Coordinators report the queries they got admitted through the
admission service in the admission heartbeat (AdmittedQueryPB: pool,
executor group, admission user, and the slots and memory still held
on each unreleased backend). QuerySchedulePB carries the pool,
executor group, trivial flag and admission user for this. To keep
heartbeats small, the full list is only sent when the admissiond asks
for it in the heartbeat response (until it has got the coordinator's
report once, i.e. after it started, or since the coordinator left its
cluster membership and had its queries released) and after a failed
heartbeat; otherwise a heartbeat only carries the ids of released
queries.
- An admissiond that does not know a reported query re-registers
("adopts") it as running via AdmissionController::AdoptRunningQuery(),
so pool and host usage are accounted for again and later
ReleaseQueryBackends/ReleaseQuery rpcs work. Released queries stay in
the report with released=true until they are unregistered, so an
adoption based on an older heartbeat is undone. New metric
admission-controller.total-adopted.<pool>.
- After it starts, the admissiond admits queries only once every
coordinator in its cluster membership has reported all its admitted
queries while in that membership, or
--admission_adoption_grace_period_ms (10 s, at least the heartbeat rpc
timeout plus two heartbeat periods) after its first admission rpc, so
that it does not admit into capacity that is still in use.
- A coordinator whose queued query is lost (GetQueryStatus fails with a
network error or times out, or returns INVALID_QUERY_HANDLE) resubmits
it with a fresh retry budget, at most 5 times, instead of failing it
(--admission_resubmit_on_admissiond_loss). The query loses its place
in the queue and its queue timeout restarts. Note that this also
applies when only the reply that rejected the query or timed it out
in the queue is lost: the admissiond removes the state of a rejected
query when it sends that reply, so the query is decided again.
- AdmitQuery (30 s) and GetQueryStatus (10 s) get rpc timeouts, and the
heartbeat gets --admission_heartbeat_rpc_timeout_ms (3 s). After a
timeout, admission rpcs use a new KRPC connection (a separate network
plane), since KRPC otherwise keeps queueing rpcs behind a connection
whose negotiation hangs until --rpc_negotiation_timeout_ms, e.g. when
the admissiond host is gone. Concurrent timeouts on one connection
switch only once. This complements IMPALA-14466, which re-resolves the
address.
Mixed versions work during a rolling upgrade: an admissiond without
this change ignores the new heartbeat fields, and does not put the pool
in the schedule, so a coordinator it admitted reports nothing. An
admissiond with this change waits for the full grace period after a
restart while some coordinators do not report yet.
Testing:
- Added AdmissionControllerTest.AdoptRunningQuery and
AdoptRunningQueryUserQuota: adoption accounts for pool, host and user
quota resources, is idempotent, and releases work as for a query
admitted locally.
- Added TestAdmissionControllerWithACService::
test_admissiond_restart_recovers_state (exhaustive only): restarts
the admissiond with one running and one queued query in a pool with
max requests 1, and checks that the running query is adopted, the
queued one is resubmitted and stays queued until the first one ends,
and no release fails. Not run yet: it needs a minicluster.
- Ran AdmissionControllerTest.*. Compiled impalad against master.
- Validated the recovery on a Kubernetes deployment, with this change
applied to a 4.5-based build, by restarting the admissiond with
queries running and queued: the running query was adopted, the queued
ones were resubmitted and ran one after another, and the pool limit
of 1 held across the restart.
Change-Id: I1c7bed8dabf880b1913190801af363dda47456f3
Assisted-by: Claude Opus 5 (Claude Code)
---
M be/src/scheduling/admission-control-client.cc
M be/src/scheduling/admission-control-client.h
M be/src/scheduling/admission-control-service.cc
M be/src/scheduling/admission-control-service.h
M be/src/scheduling/admission-controller-test.cc
M be/src/scheduling/admission-controller.cc
M be/src/scheduling/admission-controller.h
M be/src/scheduling/remote-admission-control-client.cc
M be/src/scheduling/remote-admission-control-client.h
M be/src/service/impala-server.cc
M common/protobuf/admission_control_service.proto
M common/thrift/metrics.json
M tests/custom_cluster/test_admission_controller.py
13 files changed, 959 insertions(+), 63 deletions(-)
git pull ssh://gerrit.cloudera.org:29418/Impala-ASF refs/changes/25/24925/1
--
To view, visit http://gerrit.cloudera.org:8080/24925
To unsubscribe, visit http://gerrit.cloudera.org:8080/settings
Gerrit-Project: Impala-ASF
Gerrit-Branch: master
Gerrit-MessageType: newchange
Gerrit-Change-Id: I1c7bed8dabf880b1913190801af363dda47456f3
Gerrit-Change-Number: 24925
Gerrit-PatchSet: 1
Gerrit-Owner: Anonymous Coward <[email protected]>