davidzollo opened a new pull request, #12027:
URL: https://github.com/apache/seatunnel/pull/12027

   ## Purpose
   
   Adds multi-node E2E regression coverage for the duplicate/lost pending-job
   dispatch race fixed in #11653 ("[Fix][Zeta] Avoid duplicate pending job
   scheduling after failover").
   
   That fix introduced a monotonic scheduling epoch so a scheduler thread from
   a stale master generation cannot dispatch a pending job that a newer
   generation has already claimed, and made `clearCoordinatorService`
   unconditionally drop interrupted pending jobs from the local queue so a
   later flap-back cannot re-dispatch a poisoned `JobMaster`. It already ships
   with four thorough unit tests in `CoordinatorServiceTest` that prove the
   epoch/lock mechanism correct in isolation via mocked `CoordinatorService`
   and reflection.
   
   What was still missing: proof that the same invariant holds through the
   real multi-node integration path — real client job submission, real
   Hazelcast membership changes, and real resource contention — rather than a
   single JVM driving the internal method calls directly.
   
   ## What this test does
   
   
`SplitClusterPendingJobLifecycleFailoverIT#testPendingJobNotDuplicatedAcrossRepeatedMasterFailover`:
   
   1. Starts a split-role cluster (2 master-eligible nodes + 1 worker, fixed
      4-slot capacity, `ScheduleStrategy.WAIT`).
   2. Submits a long-running "holder" job that saturates all worker slots.
   3. Submits a small bounded batch "contested" job, which goes `PENDING`
      since no slots remain. Because the strategy is `WAIT`, the scheduler
      thread repeatedly re-evaluates this job on a fixed 3-second cadence —
      this gives the test a wide, deterministic window to land repeated
      failovers while the job is under active scheduler evaluation, instead
      of chasing a microsecond-scale race.
   4. Repeatedly flaps master ownership (4 rounds): kills the active master,
      waits for the standby to take over, then starts a fresh replacement
      master-eligible node so there is always a standby ready for the next
      round.
   5. After all flaps, confirms the contested job is still tracked as a
      single `PENDING` entry (not lost, not duplicated), then cancels the
      holder job to free capacity and lets it actually run.
   6. Asserts the job reaches `FINISHED` with **exactly** the expected row
      count. The contested job is a bounded batch job specifically so this
      count is exact rather than a lower bound: if it were ever dispatched
      twice as two independent `JobMaster` instances, the sink would end up
      with double the expected rows instead of matching precisely.
   
   ## Why this belongs at the E2E layer, not just as another unit test
   
   The existing `SplitClusterPendingJobLifecycleFailoverIT` already proves a
   pending job survives a *single* clean master failover, but never times the
   kill to land while the job is actively mid-evaluation, and has no
   assertion against duplicate dispatch (only status-transition checks). The
   new test closes both gaps using entirely existing harness capabilities —
   no new test infrastructure was needed.
   
   ## Test plan
   
   - New test only; no production code changed.
   - `./mvnw spotless:apply` run on the affected module.
   - `./mvnw install -DskipTests` run on the affected module and its
     dependency chain to confirm the new test compiles cleanly (this repo's
     shaded modules need `install`, not bare `compile`/`test-compile`, to
     resolve relocated classes) — build succeeded.
   - Full test execution is left to CI per this repository's E2E conventions
     (Testcontainers/Hazelcast-backed, not run in this sandbox).


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to