[ 
https://issues.apache.org/jira/browse/FLINK-40372?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

liting liu updated FLINK-40372:
-------------------------------
    Description: 
our flink opeartor encounterd OOM after running for months.  Our flink jobs 
enable the {{periodic}} savepoint.
There are some detail logs:

 only seven {{jm-ha-f6ndtm-savepoint-periodic-*}} resources remained in the 
cluster. However, the next metrics report at {{11:18:12 CST}} still emitted 
metrics for twelve resource names. The five additional resources had already 
been deleted.
h3. Seven resources still present

Raw {{kubectl get}} output:

 

{{jm-ha-f6ndtm-savepoint-periodic-1786589833663 2026-08-13T02:57:13Z IN_PROGRESS
jm-ha-f6ndtm-savepoint-periodic-1786589864994 2026-08-13T02:57:45Z IN_PROGRESS
jm-ha-f6ndtm-savepoint-periodic-1786589927560 2026-08-13T02:58:47Z IN_PROGRESS
jm-ha-f6ndtm-savepoint-periodic-1786589943145 2026-08-13T02:59:03Z IN_PROGRESS
jm-ha-f6ndtm-savepoint-periodic-1786589958617 2026-08-13T02:59:18Z IN_PROGRESS
jm-ha-f6ndtm-savepoint-periodic-1786589974113 2026-08-13T02:59:34Z IN_PROGRESS
jm-ha-f6ndtm-savepoint-periodic-1786589983358 <none> ABANDONED}}
h3. Five additional resources had already been deleted

Directly querying these resources returned {{{}NotFound{}}}:

 

{{Error from server (NotFound): flinkstatesnapshots.flink.apache.org 
"jm-ha-f6ndtm-savepoint-periodic-1786589817875" not found
Error from server (NotFound): flinkstatesnapshots.flink.apache.org 
"jm-ha-f6ndtm-savepoint-periodic-1786589849241" not found
Error from server (NotFound): flinkstatesnapshots.flink.apache.org 
"jm-ha-f6ndtm-savepoint-periodic-1786589880549" not found
Error from server (NotFound): flinkstatesnapshots.flink.apache.org 
"jm-ha-f6ndtm-savepoint-periodic-1786589896162" not found
Error from server (NotFound): flinkstatesnapshots.flink.apache.org 
"jm-ha-f6ndtm-savepoint-periodic-1786589911865" not found}}
h3. The next reporter cycle still emitted all twelve resources

The Operator logs use UTC. Therefore, {{03:18:12}} corresponds to {{{}11:18:12 
CST{}}}.

Raw reporter header:

 

{{2026-08-13 03:18:12,173 o.a.f.m.s.Slf4jReporter [INFO ]}}

Resource metrics emitted during that report:

 

{{...jm-ha-f6ndtm-savepoint-periodic-1786589911865.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
 4
...jm-ha-f6ndtm-savepoint-periodic-1786589817875.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
 4
...jm-ha-f6ndtm-savepoint-periodic-1786589880549.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
 4
...jm-ha-f6ndtm-savepoint-periodic-1786589974113.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
 3
...jm-ha-f6ndtm-savepoint-periodic-1786589927560.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
 3
...jm-ha-f6ndtm-savepoint-periodic-1786589896162.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
 4
...jm-ha-f6ndtm-savepoint-periodic-1786589833663.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
 3
...jm-ha-f6ndtm-savepoint-periodic-1786589943145.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
 3
...jm-ha-f6ndtm-savepoint-periodic-1786589983358.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
 2
...jm-ha-f6ndtm-savepoint-periodic-1786589958617.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
 3
...jm-ha-f6ndtm-savepoint-periodic-1786589849241.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
 4
...jm-ha-f6ndtm-savepoint-periodic-1786589864994.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
 3}}

The following five entries belong to deleted resources:

 

{{1786589817875
1786589849241
1786589880549
1786589896162
1786589911865}}

The runtime sequence is therefore:

 

{{11:17 CST Only 7 matching resources existed; the other 5 returned NotFound
11:18:12 CST The next reporter cycle still emitted metrics for all 12 
resources}}

Additionally, the retained metrics for the deleted resources include:

 

{{FlinkStateSnapshot.JOSDK.Resource.Event.DELETED.Count: 1}}

This shows that the deletion events had been observed, but the corresponding 
resource-specific metrics continued to be emitted afterward.

A concise description for the issue:
{quote}Five {{jm-ha-f6ndtm-savepoint-periodic-*}} resources no longer existed 
in Kubernetes and returned {{{}NotFound{}}}. However, the subsequent SLF4J 
metrics report continued to emit their resource-specific JOSDK metrics. The 
reporter emitted metrics for 12 resource names while only 7 matching resources 
remained in the cluster. The retained metrics also contained 
{{{}Resource.Event.DELETED.Count: 1{}}}, confirming that the deletion events 
had already been observed.
{quote}

  was:
our flink opeartor encounterd OOM.  Our flink jobs enable the {{periodic}} 
savepoint.
There are some detail logs:

 only seven {{jm-ha-f6ndtm-savepoint-periodic-*}} resources remained in the 
cluster. However, the next metrics report at {{11:18:12 CST}} still emitted 
metrics for twelve resource names. The five additional resources had already 
been deleted.
h3. Seven resources still present

Raw {{kubectl get}} output:

 

{{jm-ha-f6ndtm-savepoint-periodic-1786589833663  2026-08-13T02:57:13Z  
IN_PROGRESS
jm-ha-f6ndtm-savepoint-periodic-1786589864994  2026-08-13T02:57:45Z  IN_PROGRESS
jm-ha-f6ndtm-savepoint-periodic-1786589927560  2026-08-13T02:58:47Z  IN_PROGRESS
jm-ha-f6ndtm-savepoint-periodic-1786589943145  2026-08-13T02:59:03Z  IN_PROGRESS
jm-ha-f6ndtm-savepoint-periodic-1786589958617  2026-08-13T02:59:18Z  IN_PROGRESS
jm-ha-f6ndtm-savepoint-periodic-1786589974113  2026-08-13T02:59:34Z  IN_PROGRESS
jm-ha-f6ndtm-savepoint-periodic-1786589983358  <none>                ABANDONED}}
h3. Five additional resources had already been deleted

Directly querying these resources returned {{{}NotFound{}}}:

 

{{Error from server (NotFound): flinkstatesnapshots.flink.apache.org 
"jm-ha-f6ndtm-savepoint-periodic-1786589817875" not found
Error from server (NotFound): flinkstatesnapshots.flink.apache.org 
"jm-ha-f6ndtm-savepoint-periodic-1786589849241" not found
Error from server (NotFound): flinkstatesnapshots.flink.apache.org 
"jm-ha-f6ndtm-savepoint-periodic-1786589880549" not found
Error from server (NotFound): flinkstatesnapshots.flink.apache.org 
"jm-ha-f6ndtm-savepoint-periodic-1786589896162" not found
Error from server (NotFound): flinkstatesnapshots.flink.apache.org 
"jm-ha-f6ndtm-savepoint-periodic-1786589911865" not found}}
h3. The next reporter cycle still emitted all twelve resources

The Operator logs use UTC. Therefore, {{03:18:12}} corresponds to {{{}11:18:12 
CST{}}}.

Raw reporter header:

 

{{2026-08-13 03:18:12,173 o.a.f.m.s.Slf4jReporter [INFO ]}}

Resource metrics emitted during that report:

 

{{...jm-ha-f6ndtm-savepoint-periodic-1786589911865.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
 4
...jm-ha-f6ndtm-savepoint-periodic-1786589817875.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
 4
...jm-ha-f6ndtm-savepoint-periodic-1786589880549.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
 4
...jm-ha-f6ndtm-savepoint-periodic-1786589974113.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
 3
...jm-ha-f6ndtm-savepoint-periodic-1786589927560.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
 3
...jm-ha-f6ndtm-savepoint-periodic-1786589896162.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
 4
...jm-ha-f6ndtm-savepoint-periodic-1786589833663.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
 3
...jm-ha-f6ndtm-savepoint-periodic-1786589943145.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
 3
...jm-ha-f6ndtm-savepoint-periodic-1786589983358.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
 2
...jm-ha-f6ndtm-savepoint-periodic-1786589958617.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
 3
...jm-ha-f6ndtm-savepoint-periodic-1786589849241.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
 4
...jm-ha-f6ndtm-savepoint-periodic-1786589864994.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
 3}}

The following five entries belong to deleted resources:

 

{{1786589817875
1786589849241
1786589880549
1786589896162
1786589911865}}

The runtime sequence is therefore:

 

{{11:17 CST     Only 7 matching resources existed; the other 5 returned NotFound
11:18:12 CST  The next reporter cycle still emitted metrics for all 12 
resources}}

Additionally, the retained metrics for the deleted resources include:

 

{{FlinkStateSnapshot.JOSDK.Resource.Event.DELETED.Count: 1}}

This shows that the deletion events had been observed, but the corresponding 
resource-specific metrics continued to be emitted afterward.

A concise description for the issue or PR:
{quote}Five {{jm-ha-f6ndtm-savepoint-periodic-*}} resources no longer existed 
in Kubernetes and returned {{{}NotFound{}}}. However, the subsequent SLF4J 
metrics report continued to emit their resource-specific JOSDK metrics. The 
reporter emitted metrics for 12 resource names while only 7 matching resources 
remained in the cluster. The retained metrics also contained 
{{{}Resource.Event.DELETED.Count: 1{}}}, confirming that the deletion events 
had already been observed.





{quote}


> [Kubernetes Operator] JOSDK per-resource metrics are not unregistered after 
> resource cleanup
> --------------------------------------------------------------------------------------------
>
>                 Key: FLINK-40372
>                 URL: https://issues.apache.org/jira/browse/FLINK-40372
>             Project: Flink
>          Issue Type: Bug
>          Components: Kubernetes Operator
>    Affects Versions: 1.13
>            Reporter: liting liu
>            Priority: Critical
>              Labels: pull-request-available
>
> our flink opeartor encounterd OOM after running for months.  Our flink jobs 
> enable the {{periodic}} savepoint.
> There are some detail logs:
>  only seven {{jm-ha-f6ndtm-savepoint-periodic-*}} resources remained in the 
> cluster. However, the next metrics report at {{11:18:12 CST}} still emitted 
> metrics for twelve resource names. The five additional resources had already 
> been deleted.
> h3. Seven resources still present
> Raw {{kubectl get}} output:
>  
> {{jm-ha-f6ndtm-savepoint-periodic-1786589833663 2026-08-13T02:57:13Z 
> IN_PROGRESS
> jm-ha-f6ndtm-savepoint-periodic-1786589864994 2026-08-13T02:57:45Z IN_PROGRESS
> jm-ha-f6ndtm-savepoint-periodic-1786589927560 2026-08-13T02:58:47Z IN_PROGRESS
> jm-ha-f6ndtm-savepoint-periodic-1786589943145 2026-08-13T02:59:03Z IN_PROGRESS
> jm-ha-f6ndtm-savepoint-periodic-1786589958617 2026-08-13T02:59:18Z IN_PROGRESS
> jm-ha-f6ndtm-savepoint-periodic-1786589974113 2026-08-13T02:59:34Z IN_PROGRESS
> jm-ha-f6ndtm-savepoint-periodic-1786589983358 <none> ABANDONED}}
> h3. Five additional resources had already been deleted
> Directly querying these resources returned {{{}NotFound{}}}:
>  
> {{Error from server (NotFound): flinkstatesnapshots.flink.apache.org 
> "jm-ha-f6ndtm-savepoint-periodic-1786589817875" not found
> Error from server (NotFound): flinkstatesnapshots.flink.apache.org 
> "jm-ha-f6ndtm-savepoint-periodic-1786589849241" not found
> Error from server (NotFound): flinkstatesnapshots.flink.apache.org 
> "jm-ha-f6ndtm-savepoint-periodic-1786589880549" not found
> Error from server (NotFound): flinkstatesnapshots.flink.apache.org 
> "jm-ha-f6ndtm-savepoint-periodic-1786589896162" not found
> Error from server (NotFound): flinkstatesnapshots.flink.apache.org 
> "jm-ha-f6ndtm-savepoint-periodic-1786589911865" not found}}
> h3. The next reporter cycle still emitted all twelve resources
> The Operator logs use UTC. Therefore, {{03:18:12}} corresponds to 
> {{{}11:18:12 CST{}}}.
> Raw reporter header:
>  
> {{2026-08-13 03:18:12,173 o.a.f.m.s.Slf4jReporter [INFO ]}}
> Resource metrics emitted during that report:
>  
> {{...jm-ha-f6ndtm-savepoint-periodic-1786589911865.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
>  4
> ...jm-ha-f6ndtm-savepoint-periodic-1786589817875.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
>  4
> ...jm-ha-f6ndtm-savepoint-periodic-1786589880549.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
>  4
> ...jm-ha-f6ndtm-savepoint-periodic-1786589974113.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
>  3
> ...jm-ha-f6ndtm-savepoint-periodic-1786589927560.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
>  3
> ...jm-ha-f6ndtm-savepoint-periodic-1786589896162.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
>  4
> ...jm-ha-f6ndtm-savepoint-periodic-1786589833663.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
>  3
> ...jm-ha-f6ndtm-savepoint-periodic-1786589943145.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
>  3
> ...jm-ha-f6ndtm-savepoint-periodic-1786589983358.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
>  2
> ...jm-ha-f6ndtm-savepoint-periodic-1786589958617.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
>  3
> ...jm-ha-f6ndtm-savepoint-periodic-1786589849241.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
>  4
> ...jm-ha-f6ndtm-savepoint-periodic-1786589864994.FlinkStateSnapshot.JOSDK.Resource.Event.Count:
>  3}}
> The following five entries belong to deleted resources:
>  
> {{1786589817875
> 1786589849241
> 1786589880549
> 1786589896162
> 1786589911865}}
> The runtime sequence is therefore:
>  
> {{11:17 CST Only 7 matching resources existed; the other 5 returned NotFound
> 11:18:12 CST The next reporter cycle still emitted metrics for all 12 
> resources}}
> Additionally, the retained metrics for the deleted resources include:
>  
> {{FlinkStateSnapshot.JOSDK.Resource.Event.DELETED.Count: 1}}
> This shows that the deletion events had been observed, but the corresponding 
> resource-specific metrics continued to be emitted afterward.
> A concise description for the issue:
> {quote}Five {{jm-ha-f6ndtm-savepoint-periodic-*}} resources no longer existed 
> in Kubernetes and returned {{{}NotFound{}}}. However, the subsequent SLF4J 
> metrics report continued to emit their resource-specific JOSDK metrics. The 
> reporter emitted metrics for 12 resource names while only 7 matching 
> resources remained in the cluster. The retained metrics also contained 
> {{{}Resource.Event.DELETED.Count: 1{}}}, confirming that the deletion events 
> had already been observed.
> {quote}



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to