[
https://issues.apache.org/jira/browse/FLINK-40384?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
liting liu updated FLINK-40384:
-------------------------------
Description:
{{When periodic FlinkStateSnapshot creation is enabled, active snapshot
resources may be deleted while their status is still being updated.}}
The operator repeatedly reports errors:
io.fabric8.kubernetes.client.KubernetesClientException:
Failure executing PATCH at:
.../namespaces/<namespace>/flinkstatesnapshots/<snapshot-name>
Operation cannot be fulfilled on
flinkstatesnapshots.flink.apache.org "<snapshot-name>":
the object has been modified; please apply your changes to the latest version
and try again
HTTP status: 409 Conflict
Each conflict produces a SavepointError Kubernetes Event.
The affected FlinkStateSnapshot CR is normally deleted shortly afterward, but
its Event remains until the Kubernetes Event TTL expires. Therefore, the
namespace may contain many SavepointError Events referring to
FlinkStateSnapshot resources that no longer exist.
{{How to reproduce}}
# Deploy Flink Kubernetes Operator 1.13 or later with the FlinkStateSnapshot
CRD enabled.
# {{Create a running FlinkDeployment.}}
# Configure periodic savepoints and retain only one snapshot:
kubernetes.operator.periodic.savepoint.interval: 10s
kubernetes.operator.savepoint.history.max.count: "1"
kubernetes.operator.savepoint.dispose-on-delete: "true"
then:
# {{Wait for the first periodic FlinkStateSnapshot to reach COMPLETED.}}
# {{Wait for the next periodic FlinkStateSnapshot to be created.}}
# Inspect the Kubernetes Events
was:
{{When periodic FlinkStateSnapshot creation is enabled, active snapshot
resources may be deleted while their status is still being updated.}}
The operator repeatedly reports errors:
io.fabric8.kubernetes.client.KubernetesClientException:
Failure executing PATCH at:
.../namespaces/<namespace>/flinkstatesnapshots/<snapshot-name>
Operation cannot be fulfilled on
flinkstatesnapshots.flink.apache.org "<snapshot-name>":
the object has been modified; please apply your changes to the latest version
and try again
HTTP status: 409 Conflict
Each conflict produces a SavepointError Kubernetes Event.
The affected FlinkStateSnapshot CR is normally deleted shortly afterward, but
its Event remains until the Kubernetes Event TTL expires. Therefore, the
namespace may contain many SavepointError Events referring to
FlinkStateSnapshot resources that no longer exist.
{{How to reproduce}}
#
Deploy Flink Kubernetes Operator 1.13 or later with the FlinkStateSnapshot CRD
enabled.
# {{Create a running FlinkDeployment.}}
# Configure periodic savepoints and retain only one snapshot:
kubernetes.operator.periodic.savepoint.interval: 10s
kubernetes.operator.savepoint.history.max.count: "1"
kubernetes.operator.savepoint.dispose-on-delete: "true"
then:
# {{Wait for the first periodic FlinkStateSnapshot to reach COMPLETED.}}
# {{Wait for the next periodic FlinkStateSnapshot to be created.}}
# Inspect the Kubernetes Events
> Periodic FlinkStateSnapshot cleanup causes repeated SavepointError events
> -------------------------------------------------------------------------
>
> Key: FLINK-40384
> URL: https://issues.apache.org/jira/browse/FLINK-40384
> Project: Flink
> Issue Type: Bug
> Components: Kubernetes Operator
> Reporter: liting liu
> Priority: Major
>
> {{When periodic FlinkStateSnapshot creation is enabled, active snapshot
> resources may be deleted while their status is still being updated.}}
> The operator repeatedly reports errors:
> io.fabric8.kubernetes.client.KubernetesClientException:
> Failure executing PATCH at:
> .../namespaces/<namespace>/flinkstatesnapshots/<snapshot-name>
> Operation cannot be fulfilled on
> flinkstatesnapshots.flink.apache.org "<snapshot-name>":
> the object has been modified; please apply your changes to the latest version
> and try again
> HTTP status: 409 Conflict
> Each conflict produces a SavepointError Kubernetes Event.
> The affected FlinkStateSnapshot CR is normally deleted shortly afterward, but
> its Event remains until the Kubernetes Event TTL expires. Therefore, the
> namespace may contain many SavepointError Events referring to
> FlinkStateSnapshot resources that no longer exist.
>
> {{How to reproduce}}
> # Deploy Flink Kubernetes Operator 1.13 or later with the FlinkStateSnapshot
> CRD enabled.
> # {{Create a running FlinkDeployment.}}
> # Configure periodic savepoints and retain only one snapshot:
> kubernetes.operator.periodic.savepoint.interval: 10s
> kubernetes.operator.savepoint.history.max.count: "1"
> kubernetes.operator.savepoint.dispose-on-delete: "true"
> then:
> # {{Wait for the first periodic FlinkStateSnapshot to reach COMPLETED.}}
> # {{Wait for the next periodic FlinkStateSnapshot to be created.}}
> # Inspect the Kubernetes Events
--
This message was sent by Atlassian Jira
(v8.20.10#820010)