DanielLeens commented on PR #12311:
URL: https://github.com/apache/seatunnel/pull/12311#issuecomment-5846319915

   Thanks for running this under real load — that table is exactly the evidence 
I wanted to see. 12/20 and 4/10 baseline failures collapsing to 0/20 and 0/10 
with the fix, both against the same `dev` base commit, confirms the fix closes 
the specific race the flake was hitting, and your restore/failure-reporting 
sweep across the checkpoint-error and sibling-failure paths matches my own 
trace.
   
   One thing to flag before you take this as the full picture: SEZ9 just posted 
a review on the same head that found two real edge cases your local run 
wouldn't have exercised — a master-failover-initiated cancel 
(`SubPlan#restorePipelineState`) that isn't a user cancel but now resolves 
CANCELED the same way, and a sibling vertex the sequential cancel loop hasn't 
reached yet (still RUNNING) that can still flip the job to FAILED. Both are 
timing-dependent in the same way the original flake was, so a clean 0/20 on a 
single-worker synthetic run doesn't rule them out. I'm going to fold a 
pipeline/job-status-aware version of the resolver into the next revision to 
close those too — I'll ping this thread again once that's up so you can point 
your loop at it.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to