DanielLeens commented on PR #12311: URL: https://github.com/apache/seatunnel/pull/12311#issuecomment-5846319915
Thanks for running this under real load — that table is exactly the evidence I wanted to see. 12/20 and 4/10 baseline failures collapsing to 0/20 and 0/10 with the fix, both against the same `dev` base commit, confirms the fix closes the specific race the flake was hitting, and your restore/failure-reporting sweep across the checkpoint-error and sibling-failure paths matches my own trace. One thing to flag before you take this as the full picture: SEZ9 just posted a review on the same head that found two real edge cases your local run wouldn't have exercised — a master-failover-initiated cancel (`SubPlan#restorePipelineState`) that isn't a user cancel but now resolves CANCELED the same way, and a sibling vertex the sequential cancel loop hasn't reached yet (still RUNNING) that can still flip the job to FAILED. Both are timing-dependent in the same way the original flake was, so a clean 0/20 on a single-worker synthetic run doesn't rule them out. I'm going to fold a pipeline/job-status-aware version of the resolver into the next revision to close those too — I'll ping this thread again once that's up so you can point your loop at it. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
