davidzollo commented on PR #12111:
URL: https://github.com/apache/seatunnel/pull/12111#issuecomment-5666027744

   Pushed a fix for the CI failure at 
`CheckpointCoordinatorFailoverIT#testStreamJobRecoversAfterWorkerUnreachableDuringCheckpointBarrierDispatch`
 (both JDK legs failed at `expected: <RUNNING> but was: <DEPLOYING> within 1 
minutes 8 seconds`).
   
   **The test's own premise was wrong, not the engine.** The 180s 
`hazelcast.max.no.heartbeat.seconds` override was never the blocker: 
`terminate()` on a same-host worker closes its TCP endpoint, the master's next 
connection attempt gets "Connection refused", and `MembershipManager` 
suspects/removes the member for reason "No connection" about 0.4s later -- long 
before the heartbeat ceiling or `checkpoint.timeout` could matter. Confirmed 
directly in the failing run's own log:
   
   ```
   13:27:05.093 WARN MembershipManager - Member [localhost]:5803 ... is 
suspected to be dead for reason: No connection
   ```
   
   The real cause: `SubPlan#stateProcess`'s FAILED branch calls 
`releasePipelineResource()`, then `preApplyResources(SubPlan)` (ignoring its 
boolean result), then `restorePipeline()` -> 
`ResourceUtils#applyResourceForPipeline`, which deploys whatever 
`PhysicalPlan#getPreApplyResourceFutures()` still holds -- the *original* 
submission's slot profiles. This job's pipeline needs 3 fixed slots (2 
coordinator task groups + 1 physical task group), but each worker was only 
configured with 2 (`BARRIER_DISPATCH_SLOTS_PER_WORKER = 2`). The survivor could 
never host the whole redeploy alone, so `DeployTaskOperation` kept retrying 
against the terminated member (`target=[localhost]:5803`, `invokeCount=98/100` 
in the log) using stale slot profiles, wedging the pipeline in DEPLOYING.
   
   Fix: raise `BARRIER_DISPATCH_SLOTS_PER_WORKER` to 3, add a premise guard 
that fails fast if a future template needs more slots than that, rewrite the 
Javadoc to describe the actual recovery path (connection-refused -> membership 
removal, not heartbeat/checkpoint-timeout), and record the observed `JobStatus` 
transition history in the assertion failure messages per the earlier 
non-blocking review ask. The recovery bound itself (`checkpoint.timeout + 60s`) 
is unchanged, and no other test method is touched.
   
   Separately: this same run's `SplitClusterFaultToleranceIT` failure is the 
unrelated I7 (#12030) regression being fixed in #12311 (L1 conf fix), so 
`engine-v2-it` may stay red until that lands independently of this change.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to