Rangsh commented on PR #12081: URL: https://github.com/apache/seatunnel/pull/12081#issuecomment-5994277920
@SEZ9 The re-run on `ab04db33a` is done — results for the engine-v2-it matrix: **JDK 11: fully green** (197/197 passed). **JDK 8: one error, in `ClusterFaultToleranceIT#testStreamJobRestoreInAllNodeDown`** — an awaitility `ConditionTimeoutException` waiting for the client job status to reach `CANCELED`, got `UNKNOWABLE` within 5 minutes. Run link: https://github.com/Rangsh/seatunnel/actions/runs/37257258579/job/111597488791 The key outcome either way: - **`CheckpointCoordinatorFailoverIT#testBatchJobCompletesAfterMasterFailoverDuringCloseHandshake` passed on both JDK 8 and JDK 11** (4/4 tests green on each), so the regression we traced to the `notifyCheckpointMonitor` catch swallowing `HazelcastInstanceNotActiveException` is closed by `ab04db33a`. - `BackpressureSlowSinkIT` and `SavepointBusySourceBarrierIT` also passed on both JDKs this round. - `SplitClusterFaultToleranceIT` passed (11 tests, 2 skipped) on JDK 8. For the one remaining red: `ClusterFaultToleranceIT` was green in the previous run on `1cbac0086` (both JDKs) and is green on JDK 11 here, so a single JDK-8-only occurrence looks flaky rather than related to this change — `ab04db33a` only narrows the `notifyCheckpointMonitor` catch and adds unit tests, with no code path that would delay a CANCELED transition. Per your guidance on the earlier reds, I'd treat it as suspected-flaky first and re-dispatch once on `ab04db33a` with no code changes; happy to do that now unless you'd rather inspect the logs first. I can also pull the surefire log for that test to check for anything from the WAL/checkpoint-storage path if it reproduces. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
