davidzollo commented on issue #12202: URL: https://github.com/apache/seatunnel/issues/12202#issuecomment-5665651912
Thanks @CryoThrust — agreed on both points. A timeout around `join()` before we know what is blocking would just convert a hang into a `FAILED` with the real cause lost, so I would rather not start there. I would prefer option 1: a minimal, opt-in instrumentation PR with no default behavior change — phase enter/exit and elapsed time per `cleanJob()` step (`getJobDAGInfo()`, checkpoint cleanup, history write, pending-job cleanup), plus which thread holds the `JobMaster` monitor when cleanup starts. The two candidates we could not separate from logs alone are (1) `getJobDAGInfo()`'s `synchronized(this)` contending with a concurrent `cancelJob()`/`stopJob()`, and (2) a stalled Hazelcast IMap operation in the history/cleanup writes; a trace that timestamps each step and records the monitor owner should distinguish them in a single reproduction. The fault-tolerance test that surfaced this (`SplitClusterFaultToleranceIT#testManyPipelinesRestoreContentionInWorkerDown`, #12109) needs a real multi-node cluster and only hits the window under load, so a dedicated instrumented reproduction is easier to reason about than a diagnostic flag on that test. Please reference #12202 in the PR — happy to review as soon as it is up. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
