[ 
https://issues.apache.org/jira/browse/FLINK-40384?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18104376#comment-18104376
 ] 

liting liu commented on FLINK-40384:
------------------------------------

Please assign to me

> Periodic FlinkStateSnapshot cleanup causes repeated SavepointError events
> -------------------------------------------------------------------------
>
>                 Key: FLINK-40384
>                 URL: https://issues.apache.org/jira/browse/FLINK-40384
>             Project: Flink
>          Issue Type: Bug
>          Components: Kubernetes Operator
>            Reporter: liting liu
>            Priority: Major
>
> {{When periodic FlinkStateSnapshot creation is enabled, active snapshot 
> resources may be deleted while their status is still being updated.}}
> {{The operator repeatedly reports errors similar to:
> \{code}
> io.fabric8.kubernetes.client.KubernetesClientException:
> Failure executing PATCH at:
> .../namespaces/<namespace>/flinkstatesnapshots/<snapshot-name>
> Operation cannot be fulfilled on
> flinkstatesnapshots.flink.apache.org "<snapshot-name>":
> the object has been modified; please apply your changes to the latest version 
> and try again
> HTTP status: 409 Conflict
> \{code}
> Each conflict produces a SavepointError Kubernetes Event.
> The affected FlinkStateSnapshot CR is normally deleted shortly afterward, but 
> its Event remains until the Kubernetes Event TTL expires. Therefore, the 
> namespace may contain many SavepointError Events referring to 
> FlinkStateSnapshot resources that no longer exist.}}
> {{h3. Effect}}
> {{* Every periodic savepoint cycle can generate additional unnecessary 
> SavepointError Events.
> * A large number of failure Events accumulate over time, even though only a 
> few FlinkStateSnapshot CRs currently exist.
> * Many Events refer to already-deleted resources, making the reported failure 
> impossible to inspect afterward.
> * Operator logs and Kubernetes Events are continuously polluted by 409 
> Conflict errors.
> * Genuine savepoint failures become difficult to distinguish from these 
> deletion and status-update conflicts.
> * Monitoring and alerting systems may report continuous savepoint failures 
> even when the errors are caused by the operator's internal resource lifecycle 
> handling.
> * Periodic savepoint attempts may be deleted before they complete.}}
> {{h3. How to reproduce}}
>  # {{Deploy Flink Kubernetes Operator 1.10 or later with the 
> FlinkStateSnapshot CRD enabled. }}
>  # {{Create a running FlinkDeployment.}}
>  # {{Configure periodic savepoints and retain only one snapshot:
> \{code}
> kubernetes.operator.periodic.savepoint.interval: 10s
> kubernetes.operator.savepoint.history.max.count: "1"
> kubernetes.operator.savepoint.dispose-on-delete: "true"
> \{code}}}
>  # {{Wait for the first periodic FlinkStateSnapshot to reach COMPLETED.}}
>  # {{Wait for the next periodic FlinkStateSnapshot to be created.}}
>  # {{{}{}}}{{{}Inspect the Kubernetes Events:
> \{code}
> kubectl get events -n <namespace> \
>   --field-selector reason=SavepointError \
>   --sort-by=.metadata.creationTimestamp
> \{code}{}}}
>  # {{Continue observing several periodic savepoint cycles.
> The operator repeatedly produces PATCH 409 Conflict errors and SavepointError 
> Events. After the corresponding FlinkStateSnapshot CRs are deleted, their 
> failure Events remain visible in the namespace.}}



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to