[
https://issues.apache.org/jira/browse/YUNIKORN-3381?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated YUNIKORN-3381:
-------------------------------------
Labels: pull-request-available (was: )
> Foreign pod that terminates while orphaned is later adopted as a live
> allocation
> --------------------------------------------------------------------------------
>
> Key: YUNIKORN-3381
> URL: https://issues.apache.org/jira/browse/YUNIKORN-3381
> Project: Apache YuniKorn
> Issue Type: Bug
> Components: core - scheduler, shim - kubernetes
> Reporter: Dale Richardson
> Assignee: Hedger Lai
> Priority: Major
> Labels: pull-request-available
>
> A foreign (non-YuniKorn) pod whose terminal update arrives while its node is
> unknown to the shim is never removed from the scheduler cache. When the node
> later registers, orphan adoption reports the dead pod to the core as a live
> foreign allocation, and it occupies capacity in the shim's predicate cache.
> h3. Mechanism
> {{Context.updateForeignPod}} (pkg/cache/context.go), terminated-release
> branch:
> {code}
> if oldPod != nil && utils.IsPodTerminated(pod) {
> if !ctx.schedulerCache.IsPodOrphaned(string(pod.UID)) {
> ctx.schedulerCache.RemovePod(pod) // + release to the core
> } else {
> // orphaned: "skipping foreign resource update for terminated
> orphaned pod"
> }
> return
> }
> {code}
> Skipping the release for an orphan is correct (nothing is booked in the core
> to
> release). But the branch also skips RemovePod, and it does not update the
> stored
> pod, so the cache keeps the earlier assigned, non-terminal version as an
> orphan.
> When the node arrives, {{updateNodeInternal}} (context.go:212-219) adopts the
> orphans and calls {{updateForeignPod(nil, cachedPod)}} for each foreign one;
> the
> add branch (context.go:432-441) sees an assigned, non-terminated pod and
> books a
> foreign allocation via CreateAllocationForForeignPod. The dead pod is now a
> live
> allocation in the core and an occupying pod in the node's NodeInfo.
> h3. Reproduction
> Three informer events, no scheduling required:
> # add foreign pod assigned to node N (N not yet delivered) -> orphaned, cached
> # update the same pod to Succeeded -> hits the skip branch, stale entry
> retained
> # add node N -> adoption books a foreign allocation for the
> already-terminated pod;
> the shim NodeInfo holds it as an occupying pod (pods=3 observed vs pods=2
> expected)
> Reproduced on master with a direct probe and by an informer event-ordering
> property test, which found it as a convergence divergence (delivering the same
> history with the node first books nothing).
> h3. Impact
> Phantom foreign capacity in both the core and the shim predicate cache after
> any
> window where foreign pods churn while a node is not yet (or temporarily not)
> registered -- node restarts and scheduler startup ordering are the natural
> triggers. Related to the orphan-tracking areas of YUNIKORN-2180 and
> YUNIKORN-3374,
> but a distinct defect.
> h3. Fix sketch
> In the terminated-orphaned branch, remove the pod from the cache (no core
> release
> needed): the entry has no reason to outlive the pod.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]