DanielLeens opened a new issue, #11807:
URL: https://github.com/apache/seatunnel/issues/11807

   ### Search before asking
   
   - [x] I searched existing issues and PRs and did not find a dedicated issue 
for repeated `JobHistoryService` listener registration across master role 
switches.
   
   ### What happened
   
   While reviewing the current `seatunnel-engine` source on `dev`, I found that 
`JobHistoryService` registers an IMap entry listener every time a node becomes 
the active master, but I could not find a matching listener removal path.
   
   Current source chain:
   
   1. `CoordinatorService.checkNewActiveMaster()` calls 
`initCoordinatorService()` whenever this node becomes the new active master.
   2. `initCoordinatorService()` creates a new `JobHistoryService` instance.
   3. In the `JobHistoryService` constructor, 
`finishedJobDAGInfoImap.addEntryListener(new JobInfoExpiredListener(), true)` 
is called.
   4. When the node later leaves the active-master role, 
`CoordinatorService.clearCoordinatorService()` clears executors and closes 
services, but there is no corresponding `removeEntryListener(...)` for the 
listener that was registered by the old `JobHistoryService` instance.
   
   This means repeated master failover / rolling restart can accumulate old 
listener instances. Each old listener still holds references to the old 
`JobHistoryService`, `nodeEngine`, and `logger`, and can also duplicate 
`CleanLogOperation` callbacks when historical job DAG entries expire.
   
   ### SeaTunnel Version
   
   Current `dev` branch source as of 2026-08-14.
   
   ### Reproduction / Evidence
   
   This report is based on static source analysis of the current engine code.
   
   Relevant methods/classes:
   
   - 
`org.apache.seatunnel.engine.server.CoordinatorService#checkNewActiveMaster`
   - 
`org.apache.seatunnel.engine.server.CoordinatorService#initCoordinatorService`
   - 
`org.apache.seatunnel.engine.server.CoordinatorService#clearCoordinatorService`
   - 
`org.apache.seatunnel.engine.server.master.JobHistoryService#JobHistoryService`
   - 
`org.apache.seatunnel.engine.server.master.JobHistoryService.JobInfoExpiredListener`
   
   ### Expected behavior
   
   A listener registered by `JobHistoryService` should be removed when the 
active-master coordinator is cleared, or `JobHistoryService` should expose a 
`close()` lifecycle that unregisters the listener explicitly.
   
   ### Why this matters
   
   On clusters with master switch, rolling restart, or repeated active/inactive 
transitions, listener accumulation can:
   
   - retain old service objects in heap
   - cause duplicate cleanup callbacks
   - make memory usage and expiration behavior depend on cluster history 
instead of current state only
   
   ### Possible fix direction
   
   - Store the listener registration ID returned by `addEntryListener(...)`.
   - Add an explicit `JobHistoryService.close()` or equivalent cleanup hook.
   - Call `removeEntryListener(listenerId)` from 
`CoordinatorService.clearCoordinatorService()`.
   - Add a regression test around repeated active-master transitions.
   
   ### Are you willing to submit a PR?
   
   - [ ] Yes I am willing to submit a PR!
   
   ### Code of Conduct
   
   - [x] I agree to follow this project's Code of Conduct
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to