SEZ9 commented on issue #12344:
URL: https://github.com/apache/seatunnel/issues/12344#issuecomment-5724020502

   This is a regression on `dev` itself, not a test that has always been 
broken. I can now date it, and the bisect window contains two commits that are 
suspicious on subject matter alone. This also corrects the framing in my own 
earlier comments.
   
   ### It reproduces on plain `dev`, with no pull request involved
   
   Every complete `Build` run on `apache/seatunnel@dev` since 2026-09-14 fails 
`all-connectors-it-2` on both JDKs with this exact test. Most `dev` runs are 
`cancelled` by the next merge before they finish, which is why this is easy to 
miss — but the ones that finish are unambiguous:
   
   | `dev` run | Date | Head | `OpengaussCDCIT` |
   |---|---|---|---|
   | 
[34145973469](https://github.com/apache/seatunnel/actions/runs/34145973469) | 
09-07 17:03 | `8bea8c681` | `Tests run: 23, Failures: 0, Errors: 0` in 707 s — 
**pass** |
   | 
[34752507232](https://github.com/apache/seatunnel/actions/runs/34752507232) | 
09-13 10:41 | `333913255` | *no data* — job died earlier on 
`IcebergSourceIT.startUp:161 » ContainerFetch`, never reached this class |
   | 
[34801745423](https://github.com/apache/seatunnel/actions/runs/34801745423) | 
09-14 03:11 | `46e58fc16` | `Errors: 1` — `testAddFieldWithRestore:476` in 754 
s |
   | 
[34821874806](https://github.com/apache/seatunnel/actions/runs/34821874806) | 
09-14 08:16 | `bd09f6d9a` | same failure, 799 s |
   | 
[34935118103](https://github.com/apache/seatunnel/actions/runs/34935118103) | 
09-15 06:01 | `f4a9665e8` | same failure, 786 s |
   | 
[34995029902](https://github.com/apache/seatunnel/actions/runs/34995029902) | 
09-15 16:26 | `1325a44b2` | same failure, 892 s |
   | 
[35063682504](https://github.com/apache/seatunnel/actions/runs/35063682504) | 
09-16 06:26 | `b37af3a9b` | same failure, 920 s |
   | 
[35219017572](https://github.com/apache/seatunnel/actions/runs/35219017572) | 
09-17 12:03 | `c7304ace6` | same failure, 923 s |
   
   For context, `all-connectors-it-2` did pass on `dev` on 2026-06-16, 07-29, 
08-04, 09-03 and 09-07, so this job is not chronically red — earlier red `dev` 
runs failed elsewhere (Docker Hub image fetches, for instance).
   
   **Correction to my own earlier wording.** I wrote that this test "has never 
been observed passing". That was true of every base I had measured, but all of 
those were after 09-13. It passed on `dev` a week earlier, in 707 seconds. So 
this is a regression with a date, not a permanently broken test.
   
   ### Bisect window: `8bea8c681...46e58fc16`, 81 commits
   
   The 09-13 run cannot tighten it, because that job died on an unrelated 
Docker image fetch before reaching this class. Within those 81 commits, two 
stand out because their subject lines describe precisely the mechanism this 
test exercises — add a column, savepoint, restore, expect the restored job to 
keep syncing:
   
   - **`5af8d789a` — [#11864](https://github.com/apache/seatunnel/pull/11864) 
`[Fix][Connector-V2] Fix Postgres-CDC committed-offset recovery skipping rows 
after savepoint`.** Touches `IncrementalSourceStreamFetcher`, `LsnOffset` and 
`PostgresWalFetchTask`. Opengauss CDC runs on `connector-cdc-postgres` — the 
test class is literally 
`org.apache.seatunnel.connectors.seatunnel.cdc.postgres.OpengaussCDCIT` — and 
the failing assertion is a row comparison after a savepoint and restore, which 
is the same scenario the commit title names.
   - **`ec1b1b8b5` — [#11503](https://github.com/apache/seatunnel/pull/11503) 
`[Fix][CDC][Zeta] Restore runtime schema from checkpoint after failover`.** 46 
files, +1425/-92, including `AlterTableSchemaEventHandler`, 
`TableSchemaChangeEventDispatcher`, the new `RestoreTableSchemaEvent`, 
`SeaTunnelRowDebeziumDeserializeSchema` and `IncrementalSourceReader`. Schema 
restore from a checkpoint is exactly what fails here.
   
   Three more in the window that touch the checkpoint/restore path, listed for 
completeness rather than because I have a specific reason to suspect them: 
`ff4a0b8ec` (#12150, `CheckpointCoordinator` duplicate schema-change 
completion), `dab424cd1` (#12238, `TaskExecutionService` cleanup isolation) and 
`46e58fc16` (#12260, empty barrier counters).
   
   One weak signal, offered only to be ruled out: the class's elapsed time 
grows monotonically across the failing runs — 707 s when passing, then 754, 
786, 799, 892, 920, 923. Same 23 tests each time. That may be nothing more than 
runner variance, but a restore path that has become slow enough to exceed the 
test's `atMost(60000, MILLISECONDS)` window would look exactly like this.
   
   ### What this changes about the fix
   
   My earlier offer to send an `@Disabled` quarantine PR now looks like the 
wrong instinct, and I'd rather say so than let it stand. If this is a 
regression from #11864 or #11503, the test is doing its job and silencing it 
would bury a real defect in CDC savepoint/restore — a data-correctness area. 
Reverting or fixing the responsible change is the better outcome, and 
identifying it does not require Opengauss expertise, only a bisect.
   
   I'm happy to run that bisect. The mechanism works: push a branch that is a 
target commit plus a comment-only change under `seatunnel-engine/` (enough for 
change detection to schedule `all-connectors-it-2`, invisible to spotless), 
read the one job, then cancel and delete the ref. I used it to confirm a 
different `dev`-level failure in #12353. Two or three runs across the window 
should isolate it. Say the word and I'll start, or tell me if someone closer to 
CDC would rather take it from the two candidates above — they may recognise the 
cause without any bisect at all.
   
   The quarantine option is of course still available if the community would 
rather unblock the Zeta and core-API pull requests first and fix afterwards. I 
just no longer think it should be the default.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to