abstractdog opened a new pull request, #6843:
URL: https://github.com/apache/hive/pull/6843
### What changes were proposed in this pull request?
Adds a single annotation to the Jenkins agent pod template in `Jenkinsfile`:
```yaml
metadata:
annotations:
cluster-autoscaler.kubernetes.io/safe-to-evict: "false"
```
The `yaml:` block of `hdbPodTemplate` previously contained only a `spec:`
section; this adds a
sibling `metadata:` section. The Kubernetes plugin merges this block into
the generated pod, so
no other behaviour changes — `nodeSelector`, both `type=slave` tolerations,
`fsGroup`, container
resources and volumes are untouched.
### Why are the changes needed?
Precommit builds are being killed mid-run by the GKE cluster autoscaler
deleting the node
underneath them. Jenkins reports this as:
```
Cannot contact hive-precommit-pr-NNNN-...:
hudson.remoting.ChannelClosedException:
Channel "...JNLP4-connect connection from 10.0.x.x": Remote call ...
failed.
The channel is closing down or has closed down
```
followed ~3 minutes later by `Could not connect to <pod> to send interrupt
signal to process`.
What was measured on the `hive-test-kube` cluster (project
`gcp-hive-upstream`):
* **138 distinct build pods** across ~20 PRs hit `NodeNotReady` in a 48-hour
window.
* In each case the node was deleted by a targeted
`v1.compute.instanceGroupManagers.deleteInstances` call from the cluster
autoscaler service
account (`userAgent: google-api-go-client/0.5 cluster-autoscaler`),
seconds before Jenkins lost
the channel. Two examples:
* node born 18:15:51 → deleted 18:23:34 → channel lost 18:23:38
* node born 19:23:15 → deleted 19:40:38 → channel lost 19:40:50 (this node
was running **four**
build pods for one PR)
* The nodes were healthy up to the moment of deletion — kubelet scrape
failures begin *after* the
delete call. This is not OOM, not preemption (the pool is not
spot/preemptible), not node
auto-repair, and not a Jenkins or network fault.
* Across **108 logged node removals the autoscaler recorded zero
`evictedPods`** — it does not
believe it is evicting anything.
* Node churn is severe: 372 node births / 254 removals in 48h, 94% of nodes
living under 15
minutes, and 62% of scale-down decisions reversed by a scale-up within 10
minutes.
Agent pods are bare pods (no `ownerReferences`) and carry no `safe-to-evict`
annotation, which
under documented cluster-autoscaler behaviour should block scale-down of
their node outright. It
does not here.
**Honest scope of this change:** this is a mitigation, not a verified
root-cause fix. *Why* the
autoscaler's scale-down evaluation considers these nodes removable has not
been established — GKE's
managed control plane exposes the decisions taken, not the per-node
evaluation behind them. An
initial hypothesis (pods stamped
`cloud.google.com/cluster_autoscaler_unhelpable_until: Inf` at the
pool's node ceiling, causing the autoscaler to stop counting them) was
tested and withdrawn: only 12
of the 138 killed pods were ever named in a `noScaleUp` event.
The annotation is still worth landing because it is checked on the
autoscaler's standard scale-down
path and is agnostic to *why* the node was selected. But it may not cover
the path that is actually
firing, and a GCP support case on the autoscaler behaviour should be opened
in parallel.
Two follow-ups are out of scope here, since they are cluster-side rather
than repo-side:
1. A `PodDisruptionBudget` with `maxUnavailable: 0` selecting `jenkins:
slave`. There are currently
**no PDBs anywhere** in the cluster, and node auto-upgrade is enabled
with `maxSurge: 1` and a
pending 1.34 → 1.35 node upgrade, so the next maintenance window will
drain nodes with nothing
protecting 60–110 minute builds.
2. Relieving pressure on the pool ceiling (`maxNodeCount=16`, pinned at the
ceiling ~18% of
samples). The `hdb` container requests 1800m CPU against an 8000m limit,
which caps packing at
5 pods/node.
### Does this PR introduce _any_ user-facing change?
No. CI infrastructure only — no product code, configuration or documentation
is affected.
### How was this patch tested?
No automated test is possible: this is a Jenkins pod-template annotation
whose effect is only
observable against the live GKE autoscaler.
What was verified locally:
* The `yaml:` block parses, and `metadata` and `spec` come out as sibling
top-level keys.
* The annotation value is the **string** `"false"`, not a YAML boolean — the
autoscaler requires a
string and silently ignores a boolean.
* `nodeSelector`, both tolerations and `securityContext.fsGroup` are
unchanged after the edit.
What will validate it after merge: the number of `NodeNotReady` build pods
per 48 hours, baselined
at **138** before this change. If that figure does not drop substantially,
the annotation does not
cover the autoscaler path in play and the PDB route (follow-up 1 above)
should be taken instead.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]