[
https://issues.apache.org/jira/browse/FLINK-40470?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18108133#comment-18108133
]
hyeon kim commented on FLINK-40470:
-----------------------------------
I'd like to work on this. I already have a fix with a regression test ready
locally and will open a PR once assigned.
> Savepoint redeploy loops forever when the deployment attempt fails
> ------------------------------------------------------------------
>
> Key: FLINK-40470
> URL: https://issues.apache.org/jira/browse/FLINK-40470
> Project: Flink
> Issue Type: Bug
> Components: Kubernetes Operator
> Affects Versions: kubernetes-operator-1.15.0
> Reporter: hyeon kim
> Priority: Major
>
> When {{job.savepointRedeployNonce}} is changed,
> {{AbstractJobReconciler#redeployWithSavepoint}} cancels the running job
> (deleting the whole cluster) and redeploys from {{initialSavepointPath}}.
> Unlike every other deployment path (first deployment, suspend/resume, session
> deploy), the target spec is only recorded into the resource status *after*
> the deploy call succeeds:
> {code:java}
> cancelJob(ctx, SuspendMode.STATELESS);
> ...
> deploy(...); // <- if this
> throws...
> ReconciliationUtils.updateStatusForDeployedSpec(...); // <- ...the new
> nonce is never recorded
> {code}
> If {{deploy}} throws for any reason (transient API server error, or an error
> raised after the cluster resources were actually created — e.g.
> {{KubernetesClusterDescriptor#deployApplicationCluster}} fails resolving the
> rest endpoint of a freshly created LoadBalancer service on RBAC-restricted
> clusters where listing nodes is forbidden), the new nonce is never recorded
> in {{lastReconciledSpec}}. Every subsequent reconciliation classifies the
> diff as {{SAVEPOINT_REDEPLOY}} again and repeats the full delete + redeploy
> cycle indefinitely.
> Observed in production on a managed Kubernetes platform where {{nodes}} list
> is forbidden for namespaced tenants: with
> {{kubernetes.rest-service.exposed.type: LoadBalancer}}, a savepoint redeploy
> looped delete/recreate for 23 minutes until manual intervention.
> All other job deployment paths already call
> {{ReconciliationUtils.updateStatusBeforeDeploymentAttempt}} before deploying
> (see {{AbstractJobReconciler#reconcileSpecChange}} for the suspended->running
> transition and
> {{AbstractFlinkResourceReconciler#updateStatusBeforeFirstDeployment}}), which
> makes a failed attempt recover through the regular upgrading/observe flow
> instead of retriggering the redeploy.
> *Proposed fix:* record the target spec with
> {{updateStatusBeforeDeploymentAttempt}} +
> {{statusRecorder.patchAndCacheStatus}} before the deploy call in
> {{redeployWithSavepoint}}, mirroring the suspend/resume path. A failed deploy
> then recovers on the next reconciliation by restoring from the recorded
> {{upgradeSavepointPath}} (which already points at {{initialSavepointPath}}),
> without repeating the delete + redeploy cycle.
> *Open question:* with this fix, the retry recovers through the regular
> suspended->running restore path. That path only restores from
> {{upgradeSavepointPath}} when the effective upgrade mode is not STATELESS, so
> for a spec with {{upgradeMode: stateless}} the retry falls back to an
> empty-state restore instead of the requested savepoint (still strictly better
> than the current infinite delete/redeploy loop). Distinguishing "interrupted
> savepoint redeploy" from "user deliberately switched to a stateless resume"
> in {{reconcileSpecChange}} needs a semantic decision, so I left that corner
> as a follow-up — happy to address it in this ticket if there is a preferred
> approach.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)