DanielLeens opened a new pull request, #12640: URL: https://github.com/apache/seatunnel/pull/12640
### Purpose of this pull request `unit-test (windows-latest)` is intermittently cancelled at its 120-minute limit on unrelated PRs, while a normal Windows run of that job takes about 60 minutes (median 58–62, maximum 97 across ~110 successful runs since 09-28). The job is not slow; it hangs in the teardown of `ServerExecuteCommandTest#testMemberList` after the test's assertions have already passed. Two occurrences on 2026-10-05, both with the same signature: - `unit-test (11, windows-latest)` for #12605: https://github.com/SEPURI-SAI-KRISHNA/seatunnel/actions/runs/37269092102/job/111632816285 (hung node `[localhost]:5803`) - `unit-test (11, windows-latest)` for #12152: https://github.com/CryoThrust/seatunnel/actions/runs/37306676695/job/111752918060 (hung node `[localhost]:5804`) ### Root cause (from the #12605 log) The test starts two masters and three workers in one JVM and, in `finally`, shuts them down in creation order, masters first: ``` 06:33:18.015 member table printed with 5 members (assertion passes) 06:33:18.572 [localhost]:5801 is SHUTDOWN <- active master, shut down by the test 06:33:18.572 [localhost]:5802 is SHUTTING_DOWN <- last master, shut down by the test 06:33:19.899 [localhost]:5805 All node is lite node, shutdown this cluster / Terminating forcefully... 06:33:19.919 [localhost]:5804 All node is lite node, shutdown this cluster / Terminating forcefully... 06:33:19.931 [localhost]:5803 is SHUTTING_DOWN <- graceful shutdown() from the test thread 06:33:19.933 [localhost]:5803 All node is lite node, shutdown this cluster 06:33:19.933 [localhost]:5803 Node is already shutting down... Waiting for shutdown process to complete... 06:33:22.95 5805 and 5804: Hazelcast Shutdown is completed 06:33:21 .. 07:49 [localhost]:5803 The Node is not ready yet, Node state PASSIVE (459 times, until the job is cancelled) ``` Once the last master is gone, every worker terminates itself. For 5804 and 5805 that forced termination runs alone and finishes in three seconds. For 5803 the test thread's graceful `shutdown()` wins the state transition two milliseconds earlier, so the self-termination only waits for it, and the graceful shutdown of a worker with no master left never completes. The test thread stays blocked inside `inst.shutdown()` and Surefire never gets the JVM back. Which worker loses the race, if any, is timing; on Linux runners the self-termination normally wins. ### Change Tear the cluster down in reverse creation order: the three workers first, then the standby master, then the active master. Every worker then stops while a master is still present, so the "all nodes are lite" path is never triggered during teardown and there is nothing to race. Test-only; the assertions and the cluster the test builds are unchanged. A graceful worker shutdown that can wait forever when no master exists may deserve a separate look on the engine side; this PR does not touch engine code. ### Does this PR introduce _any_ user-facing change? No. ### How was this patch tested? `./mvnw spotless:apply -pl seatunnel-core/seatunnel-starter` locally. Runtime validation is this PR's `unit-test` jobs on both Windows legs. 🤖 Generated with [Claude Code](https://claude.com/claude-code) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
