[ 
https://issues.apache.org/jira/browse/FLINK-40388?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Avi Sanwal updated FLINK-40388:
-------------------------------
    Environment: 
h4. Environment
 * Flink Kubernetes Operator: 1.15.0
 * Apache Flink: 1.20.3
 * Kubernetes: EKS

h4. Configuration
{code:yaml}
spec: 
  flinkConfiguration: 
    kubernetes.operator.job.savepoint-on-deletion: "true"
    kubernetes.operator.job.drain-on-savepoint-deletion: "true"
{code}

  was:
h4. Environment

* Flink Kubernetes Operator: 1.15.0
* Apache Flink: 1.20.3
* Kubernetes: EKS

h4. Configuration

{code:yaml}
spec:

  flinkConfiguration:

    kubernetes.operator.job.savepoint-on-deletion: "true"
    kubernetes.operator.job.drain-on-savepoint-deletion: "true"
{code}



> FlinkDeployment deletion can remain stuck when savepoint-on-deletion is 
> enabled and the job is not running
> ----------------------------------------------------------------------------------------------------------
>
>                 Key: FLINK-40388
>                 URL: https://issues.apache.org/jira/browse/FLINK-40388
>             Project: Flink
>          Issue Type: Bug
>          Components: Kubernetes Operator
>    Affects Versions: kubernetes-operator-1.15.0
>         Environment: h4. Environment
>  * Flink Kubernetes Operator: 1.15.0
>  * Apache Flink: 1.20.3
>  * Kubernetes: EKS
> h4. Configuration
> {code:yaml}
> spec: 
>   flinkConfiguration: 
>     kubernetes.operator.job.savepoint-on-deletion: "true"
>     kubernetes.operator.job.drain-on-savepoint-deletion: "true"
> {code}
>            Reporter: Avi Sanwal
>            Priority: Major
>
> h4. Problem
> Deleting a {{FlinkDeployment}} can remain in {{Terminating}} when 
> savepoint-on-deletion is enabled but the Flink job is in a non-running state, 
> such as {{RESTARTING}} during recovery or an upgrade.
> h4. Steps to reproduce
> # Deploy a stateful {{FlinkDeployment}} with high availability enabled.
> # Enable {{savepoint-on-deletion}} and {{drain-on-savepoint-deletion}}.
> # Put the job into a non-running state during recovery or an upgrade, for 
> example {{RESTARTING}}.
> # Delete the {{FlinkDeployment}} before the job reaches {{RUNNING}}.
> h4. Actual behavior
> The operator attempts deletion-time savepoint handling even though the job 
> cannot produce a savepoint. Cleanup is retried while the finalizer remains on 
> the {{FlinkDeployment}}, leaving the resource stuck in {{Terminating}} and 
> preventing a clean redeployment.
> h4. Expected behavior
> The operator should not retry an impossible savepoint operation indefinitely. 
> It should complete deletion with an explicit, observable outcome when the job 
> is not running.
> h4. Suggested resolution
> If deletion requests a savepoint but the job is not running, use a safe 
> fallback such as *last-state cleanup*:
> * Delete the Flink cluster resources.
> * Retain HA metadata so state recovery remains possible.
> * Record that the deletion-time savepoint was skipped because the job was not 
> running.
> * Remove the finalizer.
> The existing drain/savepoint behavior for a running job should remain 
> unchanged.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to