DanielLeens opened a new issue, #11806:
URL: https://github.com/apache/seatunnel/issues/11806
### Search before asking
- [x] I searched existing issues and PRs and found related work around IMAP
leaks (#9637, #9696), but I did not find an issue for the current `FAILED`
pipeline cleanup gap on the current engine source.
### What happened
While reviewing the current `seatunnel-engine` source on `dev`, I found that
pipeline metrics cleanup still skips the `FAILED` terminal path.
Current source chain:
1. `SubPlan.subPlanDone(...)` calls
`jobMaster.savePipelineMetricsToHistory(getPipelineLocation())` and then
`jobMaster.removeMetricsContext(getPipelineLocation(), pipelineStatus)`.
2. `JobMaster.removeMetricsContext(...)` only removes running metrics when:
- `pipelineStatus == FINISHED` and the pipeline is not ending by
savepoint, or
- `pipelineStatus == CANCELED`
3. `FAILED` is not included in that cleanup branch.
4. `TaskExecutionContext.getOrCreateMetricsContext(...)` reads or recreates
`SeaTunnelMetricsContext` from `IMAP_RUNNING_JOB_METRICS`, and
`SeaTunnelServer.updateMetrics(...)` keeps merging task metrics into that
distributed map.
As a result, when a pipeline ends in `FAILED`, the `TaskLocation ->
SeaTunnelMetricsContext` entries can remain in `IMAP_RUNNING_JOB_METRICS`
indefinitely.
This looks especially risky because the default
`job-metrics-partition-count` is `1`, so many leaked failed-pipeline metrics
accumulate into the same large in-memory map.
### SeaTunnel Version
Current `dev` branch source as of 2026-08-14.
### Reproduction / Evidence
This report is based on static source analysis of the current engine code.
Relevant methods:
- `org.apache.seatunnel.engine.server.dag.physical.SubPlan#subPlanDone`
- `org.apache.seatunnel.engine.server.master.JobMaster#removeMetricsContext`
- `org.apache.seatunnel.engine.server.SeaTunnelServer#updateMetrics`
- `org.apache.seatunnel.engine.server.SeaTunnelServer#removeMetrics`
-
`org.apache.seatunnel.engine.server.execution.TaskExecutionContext#getOrCreateMetricsContext`
### Expected behavior
When a pipeline reaches a terminal `FAILED` state, its running metrics
should also be removed from `IMAP_RUNNING_JOB_METRICS`, just like the
`CANCELED` path, unless there is an explicit reason to retain them elsewhere.
### Why this matters
Even after #9696 fixed earlier IMAP leak paths, this remaining
terminal-state gap can still leave failed job metrics resident in distributed
state and heap, especially under repeated failed submissions or unstable
workloads.
### Possible fix direction
- Extend `JobMaster.removeMetricsContext(...)` to cover `FAILED` terminal
pipelines.
- If there is a special case for savepoint-related failure handling,
document that boundary explicitly.
- Add an engine test that verifies `IMAP_RUNNING_JOB_METRICS` is cleaned for
`FAILED` pipelines as well as `FINISHED` and `CANCELED`.
### Are you willing to submit a PR?
- [ ] Yes I am willing to submit a PR!
### Code of Conduct
- [x] I agree to follow this project's Code of Conduct
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]