DanielLeens opened a new pull request, #12640:
URL: https://github.com/apache/seatunnel/pull/12640

   ### Purpose of this pull request
   
   `unit-test (windows-latest)` is intermittently cancelled at its 120-minute 
limit on unrelated PRs, while a normal Windows run of that job takes about 60 
minutes (median 58–62, maximum 97 across ~110 successful runs since 09-28). The 
job is not slow; it hangs in the teardown of 
`ServerExecuteCommandTest#testMemberList` after the test's assertions have 
already passed.
   
   Two occurrences on 2026-10-05, both with the same signature:
   
   - `unit-test (11, windows-latest)` for #12605: 
https://github.com/SEPURI-SAI-KRISHNA/seatunnel/actions/runs/37269092102/job/111632816285
 (hung node `[localhost]:5803`)
   - `unit-test (11, windows-latest)` for #12152: 
https://github.com/CryoThrust/seatunnel/actions/runs/37306676695/job/111752918060
 (hung node `[localhost]:5804`)
   
   ### Root cause (from the #12605 log)
   
   The test starts two masters and three workers in one JVM and, in `finally`, 
shuts them down in creation order, masters first:
   
   ```
   06:33:18.015  member table printed with 5 members (assertion passes)
   06:33:18.572  [localhost]:5801 is SHUTDOWN                      <- active 
master, shut down by the test
   06:33:18.572  [localhost]:5802 is SHUTTING_DOWN                 <- last 
master, shut down by the test
   06:33:19.899  [localhost]:5805 All node is lite node, shutdown this cluster 
/ Terminating forcefully...
   06:33:19.919  [localhost]:5804 All node is lite node, shutdown this cluster 
/ Terminating forcefully...
   06:33:19.931  [localhost]:5803 is SHUTTING_DOWN                 <- graceful 
shutdown() from the test thread
   06:33:19.933  [localhost]:5803 All node is lite node, shutdown this cluster
   06:33:19.933  [localhost]:5803 Node is already shutting down... Waiting for 
shutdown process to complete...
   06:33:22.95   5805 and 5804: Hazelcast Shutdown is completed
   06:33:21 .. 07:49  [localhost]:5803 The Node is not ready yet, Node state 
PASSIVE   (459 times, until the job is cancelled)
   ```
   
   Once the last master is gone, every worker terminates itself. For 5804 and 
5805 that forced termination runs alone and finishes in three seconds. For 5803 
the test thread's graceful `shutdown()` wins the state transition two 
milliseconds earlier, so the self-termination only waits for it, and the 
graceful shutdown of a worker with no master left never completes. The test 
thread stays blocked inside `inst.shutdown()` and Surefire never gets the JVM 
back. Which worker loses the race, if any, is timing; on Linux runners the 
self-termination normally wins.
   
   ### Change
   
   Tear the cluster down in reverse creation order: the three workers first, 
then the standby master, then the active master. Every worker then stops while 
a master is still present, so the "all nodes are lite" path is never triggered 
during teardown and there is nothing to race. Test-only; the assertions and the 
cluster the test builds are unchanged.
   
   A graceful worker shutdown that can wait forever when no master exists may 
deserve a separate look on the engine side; this PR does not touch engine code.
   
   ### Does this PR introduce _any_ user-facing change?
   
   No.
   
   ### How was this patch tested?
   
   `./mvnw spotless:apply -pl seatunnel-core/seatunnel-starter` locally. 
Runtime validation is this PR's `unit-test` jobs on both Windows legs.
   
   🤖 Generated with [Claude Code](https://claude.com/claude-code)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to