[email protected] has uploaded this change for review. ( 
http://gerrit.cloudera.org:8080/24925


Change subject: IMPALA-9976: Recover admission state after an admissiond restart
......................................................................

IMPALA-9976: Recover admission state after an admissiond restart

When the admissiond restarts it loses the state of all queries it
admitted or queued. Running queries keep running, but the new admissiond
does not count them against pool and host limits, so it over-admits
until they finish, and their later ReleaseQuery rpcs fail. Queued
queries fail once their coordinator's GetQueryStatus rpc can no longer
reach the admissiond or the new one does not know the query.

The statestore cannot be used to catch up, as the JIRA suggested: the
request-queue topic only carries per-host and per-pool aggregates, which
it deletes when the admissiond that published them fails, and the
per-query allocations needed to release a query correctly are only
known to the admissiond and to the query's coordinator. This change
makes the coordinators the source of truth:

- Coordinators report the queries they got admitted through the
  admission service in the admission heartbeat (AdmittedQueryPB: pool,
  executor group, admission user, and the slots and memory still held
  on each unreleased backend). QuerySchedulePB carries the pool,
  executor group, trivial flag and admission user for this. To keep
  heartbeats small, the full list is only sent when the admissiond asks
  for it in the heartbeat response (until it has got the coordinator's
  report once, i.e. after it started, or since the coordinator left its
  cluster membership and had its queries released) and after a failed
  heartbeat; otherwise a heartbeat only carries the ids of released
  queries.
- An admissiond that does not know a reported query re-registers
  ("adopts") it as running via AdmissionController::AdoptRunningQuery(),
  so pool and host usage are accounted for again and later
  ReleaseQueryBackends/ReleaseQuery rpcs work. Released queries stay in
  the report with released=true until they are unregistered, so an
  adoption based on an older heartbeat is undone. New metric
  admission-controller.total-adopted.<pool>.
- After it starts, the admissiond admits queries only once every
  coordinator in its cluster membership has reported all its admitted
  queries while in that membership, or
  --admission_adoption_grace_period_ms (10 s, at least the heartbeat rpc
  timeout plus two heartbeat periods) after its first admission rpc, so
  that it does not admit into capacity that is still in use.
- A coordinator whose queued query is lost (GetQueryStatus fails with a
  network error or times out, or returns INVALID_QUERY_HANDLE) resubmits
  it with a fresh retry budget, at most 5 times, instead of failing it
  (--admission_resubmit_on_admissiond_loss). The query loses its place
  in the queue and its queue timeout restarts. Note that this also
  applies when only the reply that rejected the query or timed it out
  in the queue is lost: the admissiond removes the state of a rejected
  query when it sends that reply, so the query is decided again.
- AdmitQuery (30 s) and GetQueryStatus (10 s) get rpc timeouts, and the
  heartbeat gets --admission_heartbeat_rpc_timeout_ms (3 s). After a
  timeout, admission rpcs use a new KRPC connection (a separate network
  plane), since KRPC otherwise keeps queueing rpcs behind a connection
  whose negotiation hangs until --rpc_negotiation_timeout_ms, e.g. when
  the admissiond host is gone. Concurrent timeouts on one connection
  switch only once. This complements IMPALA-14466, which re-resolves the
  address.

Mixed versions work during a rolling upgrade: an admissiond without
this change ignores the new heartbeat fields, and does not put the pool
in the schedule, so a coordinator it admitted reports nothing. An
admissiond with this change waits for the full grace period after a
restart while some coordinators do not report yet.

Testing:
 - Added AdmissionControllerTest.AdoptRunningQuery and
   AdoptRunningQueryUserQuota: adoption accounts for pool, host and user
   quota resources, is idempotent, and releases work as for a query
   admitted locally.
 - Added TestAdmissionControllerWithACService::
   test_admissiond_restart_recovers_state (exhaustive only): restarts
   the admissiond with one running and one queued query in a pool with
   max requests 1, and checks that the running query is adopted, the
   queued one is resubmitted and stays queued until the first one ends,
   and no release fails. Not run yet: it needs a minicluster.
 - Ran AdmissionControllerTest.*. Compiled impalad against master.
 - Validated the recovery on a Kubernetes deployment, with this change
   applied to a 4.5-based build, by restarting the admissiond with
   queries running and queued: the running query was adopted, the queued
   ones were resubmitted and ran one after another, and the pool limit
   of 1 held across the restart.

Change-Id: I1c7bed8dabf880b1913190801af363dda47456f3
Assisted-by: Claude Opus 5 (Claude Code)
---
M be/src/scheduling/admission-control-client.cc
M be/src/scheduling/admission-control-client.h
M be/src/scheduling/admission-control-service.cc
M be/src/scheduling/admission-control-service.h
M be/src/scheduling/admission-controller-test.cc
M be/src/scheduling/admission-controller.cc
M be/src/scheduling/admission-controller.h
M be/src/scheduling/remote-admission-control-client.cc
M be/src/scheduling/remote-admission-control-client.h
M be/src/service/impala-server.cc
M common/protobuf/admission_control_service.proto
M common/thrift/metrics.json
M tests/custom_cluster/test_admission_controller.py
13 files changed, 959 insertions(+), 63 deletions(-)



  git pull ssh://gerrit.cloudera.org:29418/Impala-ASF refs/changes/25/24925/1
--
To view, visit http://gerrit.cloudera.org:8080/24925
To unsubscribe, visit http://gerrit.cloudera.org:8080/settings

Gerrit-Project: Impala-ASF
Gerrit-Branch: master
Gerrit-MessageType: newchange
Gerrit-Change-Id: I1c7bed8dabf880b1913190801af363dda47456f3
Gerrit-Change-Number: 24925
Gerrit-PatchSet: 1
Gerrit-Owner: Anonymous Coward <[email protected]>

Reply via email to