DanielLeens commented on PR #10808:
URL: https://github.com/apache/seatunnel/pull/10808#issuecomment-5389993039
Thanks for checking, and glad the triage on the test-startup timing made
sense.
I just re-fetched my own CHANGES_REQUESTED review body directly via the
GitHub API to double-check, and it is not actually truncated on my end — the
diff block and the full "problem/root cause" explanation are all there in the
stored review. My best guess is that the long fenced diff block confused the PR
conversation view's rendering (it can happen when a code fence sits right
before a "before/after" narrative), but the underlying content is intact. For
convenience, here is that section again verbatim:
```diff
- server = createServer("server", "master");
+ // Start the lite worker first so Hazelcast mastership lands on a
worker-only member.
secondServer = createServer("secondServer", "worker");
+ server = createServer("server", "master");
```
Before this PR, the test started the non-lite `server` (master) container
first. This PR deliberately flips it to start the lite `secondServer` (worker)
container completely alone first — intentionally, per the comment, to exercise
the exact "Hazelcast mastership lands on a worker-only member" scenario this
PR's routing fix (`NodeEngineUtil.getActiveMasterAddress`) is supposed to
handle.
The problem: the pre-existing (unmodified by this PR)
`LiteNodeDropOutTcpIpJoiner` explicitly forbids a lite node from ever
self-promoting to Hazelcast master when it can't reach anyone else yet
(`LiteNodeDropOutTcpIpJoiner.joinViaPossibleMembers():157-161` and
`isThisNodeMasterCandidate():256-257`). So when `secondServer` boots alone, it
can never found the cluster by itself — the cluster is only supposed to form
once the non-lite `server` shows up. In the actual CI logs that handshake never
completes: `server` reports `Members {size:1}` once and never updates,
`secondServer` never logs `is STARTED`, and it spends the rest of the run
retrying `NodeEngineUtil.getActiveMasterAddressOrThrow` with `active master not
yet known`. That's a genuine cluster-bootstrap gap, not a routing/metrics side
effect, and it's reproduced identically across all 3 CI runs since the reorder
landed (2026-08-12, 2026-08-19, and the current head).
On the earlier-round follow-ups you listed (CLI fallback coordinator
selection vs. server's active coordinator, incompatible-changes doc coverage of
member-list semantics and the standby/failover window, telemetry doc's
bounded-timeout value/configurability, member-list indication when the active
coordinator can't be resolved, and the `java.util.Collections` import cleanup)
— those are all still open from my side and I'll re-verify them together with
Issue 3 once a fix commit lands. I don't want to re-confirm any of them as
resolved before I can see the actual diff, since the bootstrap fix itself could
touch some of that surface.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]