davidzollo commented on PR #12111: URL: https://github.com/apache/seatunnel/pull/12111#issuecomment-5666027744
Pushed a fix for the CI failure at `CheckpointCoordinatorFailoverIT#testStreamJobRecoversAfterWorkerUnreachableDuringCheckpointBarrierDispatch` (both JDK legs failed at `expected: <RUNNING> but was: <DEPLOYING> within 1 minutes 8 seconds`). **The test's own premise was wrong, not the engine.** The 180s `hazelcast.max.no.heartbeat.seconds` override was never the blocker: `terminate()` on a same-host worker closes its TCP endpoint, the master's next connection attempt gets "Connection refused", and `MembershipManager` suspects/removes the member for reason "No connection" about 0.4s later -- long before the heartbeat ceiling or `checkpoint.timeout` could matter. Confirmed directly in the failing run's own log: ``` 13:27:05.093 WARN MembershipManager - Member [localhost]:5803 ... is suspected to be dead for reason: No connection ``` The real cause: `SubPlan#stateProcess`'s FAILED branch calls `releasePipelineResource()`, then `preApplyResources(SubPlan)` (ignoring its boolean result), then `restorePipeline()` -> `ResourceUtils#applyResourceForPipeline`, which deploys whatever `PhysicalPlan#getPreApplyResourceFutures()` still holds -- the *original* submission's slot profiles. This job's pipeline needs 3 fixed slots (2 coordinator task groups + 1 physical task group), but each worker was only configured with 2 (`BARRIER_DISPATCH_SLOTS_PER_WORKER = 2`). The survivor could never host the whole redeploy alone, so `DeployTaskOperation` kept retrying against the terminated member (`target=[localhost]:5803`, `invokeCount=98/100` in the log) using stale slot profiles, wedging the pipeline in DEPLOYING. Fix: raise `BARRIER_DISPATCH_SLOTS_PER_WORKER` to 3, add a premise guard that fails fast if a future template needs more slots than that, rewrite the Javadoc to describe the actual recovery path (connection-refused -> membership removal, not heartbeat/checkpoint-timeout), and record the observed `JobStatus` transition history in the assertion failure messages per the earlier non-blocking review ask. The recovery bound itself (`checkpoint.timeout + 60s`) is unchanged, and no other test method is touched. Separately: this same run's `SplitClusterFaultToleranceIT` failure is the unrelated I7 (#12030) regression being fixed in #12311 (L1 conf fix), so `engine-v2-it` may stay red until that lands independently of this change. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
