DanielLeens commented on PR #11868:
URL: https://github.com/apache/seatunnel/pull/11868#issuecomment-5473747683

   CI record update, @JeremyXin — following up again since the `Build` check 
has been rerun to attempt 19 since my last note.
   
   That attempt (fork run `32352285042`, attempt 19, started 
2026-08-29T14:21:23Z — after my previous comment) also failed, but this time 
the concerning-looking one is `unit-test (11, windows-latest)`: 
`CheckpointCoordinatorTest > AbstractSeaTunnelServerTest.before:73 » 
IllegalState: Node failed to start!`, after a ~311.8s hang, with no 
bind-exception or other root cause captured in the surefire output.
   
   I want to be precise about why I don't think this is a regression of this 
PR's own fix, even though it's technically the same exception text and the same 
base class this PR modifies:
   - The failure took ~311s to surface (essentially a full startup-timeout 
hang), not an immediate bind/joiner failure. A genuine port-try-count-too-small 
regression would fail fast during joiner discovery, not hang for over five 
minutes.
   - I've now seen this exact signature — `IllegalStateException: Node failed 
to start!` on `windows-latest`, ~310s elapsed, a different 
`AbstractSeaTunnelServerTest` subclass each time, no captured cause — on 
another unrelated PR in this same review batch (PR #11864, `unit-test (11, 
windows-latest)` / `JobStateCleanupDelayTest`, ~311s). Two independent PRs 
touching completely different code (Postgres-CDC vs. Hazelcast join config) 
hitting the identical timing/exception fingerprint on the same OS/runner is 
strong evidence this is a shared Windows CI infra flake, not something either 
diff introduced.
   - The actual fix this PR ships (`hazelcast.tcp.join.port.try.count` raised 
to match `port-count: 100`) is still present and correct in the current head, 
and I re-confirmed the previously-failing `OptionRulesApiTest` Windows job (the 
concrete regression I originally blocked on) has not recurred in this or the 
prior attempt.
   
   My "Ready to merge" conclusion from 2026-08-20 stands. At 19 rerun attempts 
with a different unrelated flaky job/class failing almost every time, I don't 
think further reruns are a productive way to reach a clean run — I'd recommend 
a maintainer with write access merge on the strength of the code review rather 
than continuing to wait for an all-green attempt.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to