[
https://issues.apache.org/jira/browse/YUNIKORN-3382?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated YUNIKORN-3382:
-------------------------------------
Labels: pull-request-available (was: )
> Failed victim pod delete permanently wedges a preemption; release-path
> DeleteTaskPod is never retried
> -----------------------------------------------------------------------------------------------------
>
> Key: YUNIKORN-3382
> URL: https://issues.apache.org/jira/browse/YUNIKORN-3382
> Project: Apache YuniKorn
> Issue Type: Bug
> Components: core - scheduler, shim - kubernetes
> Reporter: Dale Richardson
> Assignee: Yen-Hua Chen
> Priority: Major
> Labels: pull-request-available
>
> When the shim executes a core-initiated release by deleting the pod, a failed
> delete is logged and dropped -- there is no retry on this path. One transient
> apiserver failure leaves the release permanently half-executed.
> h3. Mechanism
> The core marks a preemption victim and hands it to the RM to release, but
> does not
> free its capacity until the shim confirms:
> - victim.MarkPreempted() at preempt time (preemption.go:659), before the RM is
> notified -- so the victim is excluded from all future victim selection
> (queue.go:2206) from that moment;
> - victimQueue.IncPreemptingResource(...) (preemption.go:681);
> - p.ask.MarkTriggeredPreemption() (preemption.go:702);
> - notifyRMAllocationReleased(PREEMPTED_BY_SCHEDULER) (preemption.go:705).
> The victim's node/queue release and DecPreemptingResource run only when the
> shim's
> PREEMPTED_BY_SCHEDULER confirmation returns (partition.go:1562-1563, gated by
> the
> "this release is the shim's confirmation" logic at partition.go:1577-1579).
> The shim executes the release by deleting the pod, and drops any failure:
> {code}
> // Application.handleReleaseAppAllocationEvent (application.go:648-655)
> err := task.DeleteTaskPod()
> if err != nil {
> log.Log(...).Error("failed to release allocation from application",
> zap.Error(err))
> }
> {code}
> Nothing re-queues the task, no timer retries the delete, and the core is never
> told the release did not happen. MarkUnPreempted only runs on the "victim
> already
> released" revert (preemption.go:669), never on the failed-delete path. The
> result
> is permanent:
> - the victim keeps running and stays IsPreempted forever, excluded from future
> preemption;
> - the victim queue's preemptingResource is never decremented, so it
> under-reports
> its preemptable and remaining-guaranteed headroom by the victim's size;
> - the high-priority ask the preemption was made for was marked as having
> triggered
> preemption and now starves -- its capacity never arrives.
> h3. Reproduction
> Chaos scenario: two-queue guaranteed/max config, cluster filled with
> low-priority
> pods, high-priority wave forces preemption, victim deletes failed with p=0.5.
> After settling, the victim runs on both sides consistently (no accounting
> divergence), the queue holds a non-zero preemptingResource forever, and one
> high-priority pod is never placed. One failed delete per wedge; deterministic
> per
> seed. The same dropped-delete leaves released gang placeholders alive
> indefinitely.
> h3. Impact
> Transient apiserver delete failures (timeouts, webhook hiccups) are routine in
> busy clusters. Each occurrence silently degrades preemption capacity and
> guaranteed-quota enforcement for the victim queue until a scheduler restart,
> and
> starves the ask that triggered the preemption.
> h3. Fix sketch
> Retry the delete (bounded backoff) or re-drive it from a reconciliation pass,
> or
> report the failed release back to the core so it can re-issue. The single
> release
> path above covers both the preemption and the gang-placeholder consequence.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]