DanielLeens commented on PR #11868: URL: https://github.com/apache/seatunnel/pull/11868#issuecomment-5473747683
CI record update, @JeremyXin — following up again since the `Build` check has been rerun to attempt 19 since my last note. That attempt (fork run `32352285042`, attempt 19, started 2026-08-29T14:21:23Z — after my previous comment) also failed, but this time the concerning-looking one is `unit-test (11, windows-latest)`: `CheckpointCoordinatorTest > AbstractSeaTunnelServerTest.before:73 » IllegalState: Node failed to start!`, after a ~311.8s hang, with no bind-exception or other root cause captured in the surefire output. I want to be precise about why I don't think this is a regression of this PR's own fix, even though it's technically the same exception text and the same base class this PR modifies: - The failure took ~311s to surface (essentially a full startup-timeout hang), not an immediate bind/joiner failure. A genuine port-try-count-too-small regression would fail fast during joiner discovery, not hang for over five minutes. - I've now seen this exact signature — `IllegalStateException: Node failed to start!` on `windows-latest`, ~310s elapsed, a different `AbstractSeaTunnelServerTest` subclass each time, no captured cause — on another unrelated PR in this same review batch (PR #11864, `unit-test (11, windows-latest)` / `JobStateCleanupDelayTest`, ~311s). Two independent PRs touching completely different code (Postgres-CDC vs. Hazelcast join config) hitting the identical timing/exception fingerprint on the same OS/runner is strong evidence this is a shared Windows CI infra flake, not something either diff introduced. - The actual fix this PR ships (`hazelcast.tcp.join.port.try.count` raised to match `port-count: 100`) is still present and correct in the current head, and I re-confirmed the previously-failing `OptionRulesApiTest` Windows job (the concrete regression I originally blocked on) has not recurred in this or the prior attempt. My "Ready to merge" conclusion from 2026-08-20 stands. At 19 rerun attempts with a different unrelated flaky job/class failing almost every time, I don't think further reruns are a productive way to reach a clean run — I'd recommend a maintainer with write access merge on the strength of the code review rather than continuing to wait for an all-green attempt. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
