davidzollo commented on issue #12202:
URL: https://github.com/apache/seatunnel/issues/12202#issuecomment-5665651912

   Thanks @CryoThrust — agreed on both points. A timeout around `join()` before 
we know what is blocking would just convert a hang into a `FAILED` with the 
real cause lost, so I would rather not start there.
   
   I would prefer option 1: a minimal, opt-in instrumentation PR with no 
default behavior change — phase enter/exit and elapsed time per `cleanJob()` 
step (`getJobDAGInfo()`, checkpoint cleanup, history write, pending-job 
cleanup), plus which thread holds the `JobMaster` monitor when cleanup starts. 
The two candidates we could not separate from logs alone are (1) 
`getJobDAGInfo()`'s `synchronized(this)` contending with a concurrent 
`cancelJob()`/`stopJob()`, and (2) a stalled Hazelcast IMap operation in the 
history/cleanup writes; a trace that timestamps each step and records the 
monitor owner should distinguish them in a single reproduction.
   
   The fault-tolerance test that surfaced this 
(`SplitClusterFaultToleranceIT#testManyPipelinesRestoreContentionInWorkerDown`, 
#12109) needs a real multi-node cluster and only hits the window under load, so 
a dedicated instrumented reproduction is easier to reason about than a 
diagnostic flag on that test.
   
   Please reference #12202 in the PR — happy to review as soon as it is up.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to