SEZ9 commented on PR #12494:
URL: https://github.com/apache/seatunnel/pull/12494#issuecomment-5923907202

   Thanks for offering to take the reproduction/evidence side, @yigitcan-ozturk 
— that lines up well with what is still open on this PR, and using this thread 
as the coordination point works for me.
   
   The sequence you listed is exactly the one behind the two HIGH findings 
here: a reset-triggered (or delayed, old-generation) `CANCELED` notification 
carries only the `TaskGroupLocation`, so the master attributes it to whatever 
execution is currently at that location. The expected failure signature is the 
freshly deployed vertex going `RUNNING→CANCELED` and its slot being released 
while the new task group keeps running on the worker, possibly followed by a 
spurious pipeline restore. A few concrete asks so the evidence maps cleanly 
onto those findings:
   
   1. For each `CANCELED` notification observed by the master, please record 
which execution context (old vs. new generation) actually emitted it, and which 
execution the master applied it to. That is the core of F1/F2 and the one thing 
the current logs cannot distinguish.
   2. Please include the case where the old generation is still active on the 
worker when the replacement is deployed at the same `TaskGroupLocation` (not 
only the case where it has already exited). The existing unit test does not 
deterministically hit the "is being reset" deploy branch because `BlockTask` 
exits on interrupt before the redeploy, so a scenario that holds the old 
context alive across the redeploy would be very useful.
   3. For the IT side: the current `SplitClusterFaultToleranceIT` resets 
workers that never actually left the cluster while the master still tracks them 
as `RUNNING`, so it behaves more like a failover test than a merge. If you can 
drive a real eviction-after-heartbeat-timeout followed by rejoin/merge, that 
would give us the topology this PR targets.
   4. If you have the chance, note where `reset()` runs (merge thread) and 
whether it blocks behind an in-flight `deployTask` (`taskGroup.init()` / 
`task.init()` / `startedLatch.await()`), since that stall is the MEDIUM 
robustness finding.
   5. If any of your nodes run as `MASTER_AND_WORKER`, please also capture 
whether the losing side's coordinator/slot state survives the merge — 
`SeaTunnelServer.reset()` currently only resets the worker side.
   
   Member identities, task-group locations, slot ownership and checkpoint 
progression, as you proposed, are all welcome. Please post the raw sequence 
here when you have it; once we can see the attribution path concretely, we can 
settle on how the notification should carry execution identity.
   
   <!-- streview-comment:1445 -->


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to