[
https://issues.apache.org/jira/browse/FLINK-40226?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18098564#comment-18098564
]
Aniruddh J commented on FLINK-40226:
------------------------------------
If this looks like a valid case, I would like to fix it.
> Standalone session cluster JM recovery enters infinite MISSING loop and fails
> with HTTP 409 on TM Deployment when TaskManager pods still exist
> ----------------------------------------------------------------------------------------------------------------------------------------------
>
> Key: FLINK-40226
> URL: https://issues.apache.org/jira/browse/FLINK-40226
> Project: Flink
> Issue Type: Bug
> Components: Kubernetes Operator
> Affects Versions: 1.11.0, 1.14.0
> Environment: Operator version 1.11.0, 1.14.0 (the ones I have tested
> it on)
> Flink version1.20.x (standalone mode)
> Cluster typeStandalone Kubernetes session cluster ({{{}FlinkDeployment{}}}
> with no {{{}spec.job{}}})
> Kubernetes reproduced on OpenShift and vanilla Kubernetes
> HA Enabled
> Reporter: Aniruddh J
> Priority: Major
>
> When the *JobManager (JM) Deployment* of a standalone Kubernetes session
> cluster is deleted externally (e.g. by an admin) while the {*}TaskManager
> (TM) Deployment and its pods are still running or terminating{*}, the
> operator enters a state where it can never successfully recover the session
> cluster. Two distinct but related defects combine to produce this outcome:
> # *Bug 1 — Infinite MISSING status loop on submit failure.*
> {{recoverSession()}} in {{SessionReconciler}} does not set the
> {{JobManagerDeploymentStatus}} to {{DEPLOYING}} before calling
> {{{}submitSessionCluster(){}}}. If the submit call throws, the controller's
> error-handling path resets the status back to {{{}MISSING{}}}, causing
> {{shouldRecoverDeployment()}} to fire on every subsequent reconcile cycle —
> producing an unbound retry loop with no back-off.
> # *Bug 2 — HTTP 409 AlreadyExists on TM Deployment re-creation.*
> {{deployClusterInternal()}} in {{KubernetesStandaloneClusterDescriptor}} uses
> a bare Kubernetes {{CREATE}} (HTTP POST) for the TaskManager Deployment via
> {{{}Fabric8FlinkStandaloneKubeClient.createTaskManagerDeployment(){}}}. If
> the TM Deployment already exists — which is the normal state when only the JM
> was deleted — the API server returns {{{}HTTP 409 Conflict
> (AlreadyExists){}}}, causing every recovery attempt to fail. This failure
> feeds back into Bug 1.
>
> Together, the operator logs a continuous stream of reconciliation errors and
> the session cluster remains permanently unrecoverable until manual operator
> intervention.
>
>
> h2. Steps to Reproduce
> # Deploy a standalone Kubernetes session cluster (HA mode enabled) using
> {{{}FlinkDeployment{}}}. Wait for {{{}jobManagerDeploymentStatus: READY{}}}.
> # Manually delete only the *JM Deployment* from the namespace:
> {{kubectl delete deployment <cluster-id> -n <namespace>}}
> # Do _not_ delete the TM Deployment or its pods — allow them to continue
> running.
> # Observe the operator logs. The operator detects {{MISSING}} and calls
> {{{}recoverSession(){}}}.
> # {{deployClusterInternal()}} attempts to re-create the TM Deployment via a
> bare HTTP POST. Because the TM Deployment still exists, the API server
> returns {{{}HTTP 409 AlreadyExists{}}}.
> # The 409 exception propagates through {{submitSessionCluster()}} →
> {{recoverSession()}} → controller error handler.
> # The controller error handler resets the status to {{MISSING}} from the
> in-memory cache (which was never updated to {{{}DEPLOYING{}}}).
> # On the next reconcile cycle (15 s default), {{shouldRecoverDeployment()}}
> evaluates to {{true}} again, and the loop repeats indefinitely.
>
> h2. Expected Behaviour
> # When a JM-recovery submit fails for any transient reason (network error,
> 409, throttling), the {{jobManagerDeploymentStatus}} should be advanced to
> {{DEPLOYING}} _before_ the submit attempt, so that the error path does not
> restore {{MISSING}} and re-trigger an immediate retry.
> # A re-submit when only the JM Deployment is absent and the TM Deployment
> still exists should succeed without a 409. The desired TM spec is unchanged,
> so an idempotent create-or-replace is semantically correct.
> # A failed recovery attempt should not proactively delete or disturb
> still-running TM pods as part of its error rollback.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)