SEZ9 commented on PR #12494: URL: https://github.com/apache/seatunnel/pull/12494#issuecomment-5923907202
Thanks for offering to take the reproduction/evidence side, @yigitcan-ozturk — that lines up well with what is still open on this PR, and using this thread as the coordination point works for me. The sequence you listed is exactly the one behind the two HIGH findings here: a reset-triggered (or delayed, old-generation) `CANCELED` notification carries only the `TaskGroupLocation`, so the master attributes it to whatever execution is currently at that location. The expected failure signature is the freshly deployed vertex going `RUNNING→CANCELED` and its slot being released while the new task group keeps running on the worker, possibly followed by a spurious pipeline restore. A few concrete asks so the evidence maps cleanly onto those findings: 1. For each `CANCELED` notification observed by the master, please record which execution context (old vs. new generation) actually emitted it, and which execution the master applied it to. That is the core of F1/F2 and the one thing the current logs cannot distinguish. 2. Please include the case where the old generation is still active on the worker when the replacement is deployed at the same `TaskGroupLocation` (not only the case where it has already exited). The existing unit test does not deterministically hit the "is being reset" deploy branch because `BlockTask` exits on interrupt before the redeploy, so a scenario that holds the old context alive across the redeploy would be very useful. 3. For the IT side: the current `SplitClusterFaultToleranceIT` resets workers that never actually left the cluster while the master still tracks them as `RUNNING`, so it behaves more like a failover test than a merge. If you can drive a real eviction-after-heartbeat-timeout followed by rejoin/merge, that would give us the topology this PR targets. 4. If you have the chance, note where `reset()` runs (merge thread) and whether it blocks behind an in-flight `deployTask` (`taskGroup.init()` / `task.init()` / `startedLatch.await()`), since that stall is the MEDIUM robustness finding. 5. If any of your nodes run as `MASTER_AND_WORKER`, please also capture whether the losing side's coordinator/slot state survives the merge — `SeaTunnelServer.reset()` currently only resets the worker side. Member identities, task-group locations, slot ownership and checkpoint progression, as you proposed, are all welcome. Please post the raw sequence here when you have it; once we can see the attribution path concretely, we can settle on how the notification should carry execution identity. <!-- streview-comment:1445 --> -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
