abstractdog opened a new pull request, #6843:
URL: https://github.com/apache/hive/pull/6843

   ### What changes were proposed in this pull request?
   
   Adds a single annotation to the Jenkins agent pod template in `Jenkinsfile`:
   
   ```yaml
   metadata:
     annotations:
       cluster-autoscaler.kubernetes.io/safe-to-evict: "false"
   ```
   
   The `yaml:` block of `hdbPodTemplate` previously contained only a `spec:` 
section; this adds a
   sibling `metadata:` section. The Kubernetes plugin merges this block into 
the generated pod, so
   no other behaviour changes — `nodeSelector`, both `type=slave` tolerations, 
`fsGroup`, container
   resources and volumes are untouched.
   
   ### Why are the changes needed?
   
   Precommit builds are being killed mid-run by the GKE cluster autoscaler 
deleting the node
   underneath them. Jenkins reports this as:
   
   ```
   Cannot contact hive-precommit-pr-NNNN-...: 
hudson.remoting.ChannelClosedException:
     Channel "...JNLP4-connect connection from 10.0.x.x": Remote call ... 
failed.
     The channel is closing down or has closed down
   ```
   
   followed ~3 minutes later by `Could not connect to <pod> to send interrupt 
signal to process`.
   
   What was measured on the `hive-test-kube` cluster (project 
`gcp-hive-upstream`):
   
   * **138 distinct build pods** across ~20 PRs hit `NodeNotReady` in a 48-hour 
window.
   * In each case the node was deleted by a targeted
     `v1.compute.instanceGroupManagers.deleteInstances` call from the cluster 
autoscaler service
     account (`userAgent: google-api-go-client/0.5 cluster-autoscaler`), 
seconds before Jenkins lost
     the channel. Two examples:
     * node born 18:15:51 → deleted 18:23:34 → channel lost 18:23:38
     * node born 19:23:15 → deleted 19:40:38 → channel lost 19:40:50 (this node 
was running **four**
       build pods for one PR)
   * The nodes were healthy up to the moment of deletion — kubelet scrape 
failures begin *after* the
     delete call. This is not OOM, not preemption (the pool is not 
spot/preemptible), not node
     auto-repair, and not a Jenkins or network fault.
   * Across **108 logged node removals the autoscaler recorded zero 
`evictedPods`** — it does not
     believe it is evicting anything.
   * Node churn is severe: 372 node births / 254 removals in 48h, 94% of nodes 
living under 15
     minutes, and 62% of scale-down decisions reversed by a scale-up within 10 
minutes.
   
   Agent pods are bare pods (no `ownerReferences`) and carry no `safe-to-evict` 
annotation, which
   under documented cluster-autoscaler behaviour should block scale-down of 
their node outright. It
   does not here.
   
   **Honest scope of this change:** this is a mitigation, not a verified 
root-cause fix. *Why* the
   autoscaler's scale-down evaluation considers these nodes removable has not 
been established — GKE's
   managed control plane exposes the decisions taken, not the per-node 
evaluation behind them. An
   initial hypothesis (pods stamped 
`cloud.google.com/cluster_autoscaler_unhelpable_until: Inf` at the
   pool's node ceiling, causing the autoscaler to stop counting them) was 
tested and withdrawn: only 12
   of the 138 killed pods were ever named in a `noScaleUp` event.
   
   The annotation is still worth landing because it is checked on the 
autoscaler's standard scale-down
   path and is agnostic to *why* the node was selected. But it may not cover 
the path that is actually
   firing, and a GCP support case on the autoscaler behaviour should be opened 
in parallel.
   
   Two follow-ups are out of scope here, since they are cluster-side rather 
than repo-side:
   
   1. A `PodDisruptionBudget` with `maxUnavailable: 0` selecting `jenkins: 
slave`. There are currently
      **no PDBs anywhere** in the cluster, and node auto-upgrade is enabled 
with `maxSurge: 1` and a
      pending 1.34 → 1.35 node upgrade, so the next maintenance window will 
drain nodes with nothing
      protecting 60–110 minute builds.
   2. Relieving pressure on the pool ceiling (`maxNodeCount=16`, pinned at the 
ceiling ~18% of
      samples). The `hdb` container requests 1800m CPU against an 8000m limit, 
which caps packing at
      5 pods/node.
   
   ### Does this PR introduce _any_ user-facing change?
   
   No. CI infrastructure only — no product code, configuration or documentation 
is affected.
   
   ### How was this patch tested?
   
   No automated test is possible: this is a Jenkins pod-template annotation 
whose effect is only
   observable against the live GKE autoscaler.
   
   What was verified locally:
   
   * The `yaml:` block parses, and `metadata` and `spec` come out as sibling 
top-level keys.
   * The annotation value is the **string** `"false"`, not a YAML boolean — the 
autoscaler requires a
     string and silently ignores a boolean.
   * `nodeSelector`, both tolerations and `securityContext.fsGroup` are 
unchanged after the edit.
   
   What will validate it after merge: the number of `NodeNotReady` build pods 
per 48 hours, baselined
   at **138** before this change. If that figure does not drop substantially, 
the annotation does not
   cover the autoscaler path in play and the PDB route (follow-up 1 above) 
should be taken instead.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to