SEZ9 commented on PR #12298:
URL: https://github.com/apache/seatunnel/pull/12298#issuecomment-5707445475

   ## CI status after four more rerun rounds
   
   I reran the failed jobs four more times (attempts 5–8). One job cleared, one 
did not.
   
   **Final state:** `all-connectors-it-2 (8, 11)` only.
   
   ### `all-connectors-it-1 (11)` — cleared, and my earlier numbers for it were 
wrong
   
   `NebulaGraphIT` passed on attempt 6. But while re-checking I found a 
methodology error that affects everything I posted about this job, so the 
retraction matters more than the pass.
   
   When `rerun --failed` creates a new attempt, GitHub carries the jobs it did 
*not* re-execute forward into that attempt, keeping their old conclusion **and 
an identical `started_at`**. I read those as independent executions. They are 
not. Deduplicating by distinct `started_at`, the real counts for 
`all-connectors-it-1` are:
   
   | Base | JDK 8 | JDK 11 |
   |---|---|---|
   | `75b60fa14` — this PR | 1 real execution, 0 failures | 6 real, 5 failures 
(a1–a5 fail, a6 pass) |
   | `75b60fa14` — #12299 | 1 real execution, 0 failures | 3 real, 2 failures |
   | `0d9f9e230` (pre-#11201) — this PR | 1 real, 0 failures | 1 real, 0 
failures |
   
   So my earlier "JDK 8 passes 8/8" and "0/8 before vs 6/8 after" were both 
inflated: JDK 8 only ever executed the job **once per run**, because it passed 
first try and `--failed` never scheduled it again. The JDK asymmetry rests on 2 
samples, and the pre-#11201 side on 1 per JDK — far too thin to locate a 
regression window. I have posted this correction on #12345 and downgraded the 
`${testcontainer.version}` lead there to a guess. What survives is only that 
the JDK 11 leg failed 7 of 9 real executions on this base, which is enough to 
make the job unreliable but says nothing about cause.
   
   ### `all-connectors-it-2 (8, 11)` — #12344
   
   `OpengaussCDCIT.testAddFieldWithRestore`: **40 failures in 40 real 
executions**, across this PR and #12299, two bases, both JDKs, no pass ever 
observed. Tracked as #12344. This one is unaffected by the carry-forward error 
— it failed on every attempt, so every attempt really re-ran it, and the 
`started_at` values are all distinct.
   
   ### An observation this PR contributes to a separate dev bug
   
   While investigating #12299's `engine-v2-it (8)` failure I used this PR as a 
control, which turned up something worth recording. 
`SplitClusterFaultToleranceIT.testStreamJobCancelResolvesWhenWorkerCrashesBeforeCancelAck`
 (upstream #12353, `expected: <CANCELED> but was: <FAILED>`) passed on all 
**3** real executions of `engine-v2-it (8)` here, while failing **8/8** on 
#12299 and also failing on a control branch that is this same base plus nothing 
but a README comment. So the bug is base-level, but this PR appears to avoid 
it. The only difference I can point at is suite composition — this PR runs 204 
tests in that module because it adds 
`RestApiIT.testDynamicHttpPortIsResolvableByPeers`, while the base and #12299 
run 203 and fail. For a test built around a deliberate race between a cancel 
request and a worker shutdown, a timing shift is plausible, but I have not 
tested it and am not claiming it. Noting it in case it helps whoever picks up 
#12353.
   
   ### This PR's own change
   
   `RestApiIT`: 23 run, 0 failures, **0 skipped** on both JDK 8 and JDK 11 — 
the `Skipped: 0` is what rules out the new 
`testDynamicHttpPortIsResolvableByPeers` being silently skipped rather than 
actually passing.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to