DanielLeens commented on PR #12130: URL: https://github.com/apache/seatunnel/pull/12130#issuecomment-5605329874
Follow-up on the CI side, since a clean `Build` run is the only thing still open here. I pulled the job logs rather than going by the run summary. **What is actually red.** On fork run [`awsomesud347/seatunnel` 34089819303](https://github.com/awsomesud347/seatunnel/actions/runs/34089819303) the only job that never went green across attempts 1-4 is `kudu-connector-it (11)`, cancelled at the 90-minute job limit every time. The three jobs that failed in attempt 1 (`all-connectors-it-4 (11)`, `paimon-connector-it (8)`, `rocketmq-connector-it (8)`) all passed on re-run with no code change, and `kudu-connector-it (8)` passed in attempt 1 in 27 minutes. - attempt 1, [job 101641543194](https://github.com/awsomesud347/seatunnel/actions/runs/34089819303/job/101641543194): `KuduIT#testKuduMultipleReadWithRegex` on the Flink 1.15.3 leg. The TaskManager logs `The heartbeat of JobManager ... timed out` at 06:48:06, then `Could not resolve JobManager address ...` every 10 s until it exits at 06:53:21 with `RegistrationTimeoutException`; after that the log is silent for 71 minutes because the test JVM is still inside `container.executeJob`. - attempt 4, [job 101844792732](https://github.com/awsomesud347/seatunnel/actions/runs/34089819303/job/101844792732): `KuduIT#testKuduMultipleRead` on the Flink 1.18.0 leg. The JobMaster logs `Deploying Source: Kudu-Source -> MultiTableSink-Sink: Writer (1/1) (attempt #0)` at 19:39:48 and never logs again. The TaskManager logs `BlobClient - Downloading .../p-... from jobmanager` at 19:39:48, `Slot offering to JobManager did not finish in time` at 19:39:58 and `The heartbeat of JobManager ... timed out` at 19:41:48. From there until the cancel at 20:47 the ResourceManager/TaskManager slot-allocation loop just repeats (446 `DefaultSlotStatusSyncer` and 753 `DefaultJobLeaderService` lines) and `flink run` never returns. That is the signature tracked in #12132: the JobMaster stops answering at the first task deployment of the first job on a freshly started JobManager/TaskManager pair, on the 1.15.3 and 1.18.0 legs in PR runs (and on the 1.16.0 and 1.17.2 legs in the scheduled full matrix), never so far on 1.13.6 or 1.20.1, and `AbstractTestFlinkContainer.executeJob` has no upper bound, so the job burns the whole budget. It is not specific to this branch, this commit or JDK 11: - the same hang hit the **JDK 8** leg of this PR's earlier run ([run 34001671707 attempt 1, job 101401900552](https://github.com/awsomesud347/seatunnel/actions/runs/34001671707/job/101401900552): Flink 1.18.0, `testKuduWholeDatabaseRead`, deploy at 01:23:18, heartbeat timeout at 01:25:18) while the JDK 11 leg passed that time; - on 2026-09-09 it took both Kudu legs of #12203 ([DanielLeens/seatunnel run 34207958769](https://github.com/DanielLeens/seatunnel/actions/runs/34207958769), same deploy-then-heartbeat-timeout trace on Flink 1.18.0), apache's own scheduled dev run [34044345096](https://github.com/apache/seatunnel/actions/runs/34044345096) lost both Kudu legs the same way on 2026-09-06 (Flink 1.16.0 and 1.17.2 legs, deploy then heartbeat timeout 120 s later), and #12169 and #12165 each lost one Kudu leg the same way. Nothing in this diff can reach that path: the change is confined to `PageBaseServlet`, `JobInfoService` and `FinishedJobsServlet` in `seatunnel-engine-server`, and none of `connector-kudu`, `seatunnel-e2e-common`, `connector-kudu-e2e` or the Flink starters declare a dependency on that module; the Flink e2e containers only get the Flink starter and the connector jars. **Two things to move this along:** 1. @awsomesud347 only the fork owner can re-run a job on your fork. Could you re-run *just the failed job* on run 34089819303 (Actions, open the run, "Re-run failed jobs") rather than pushing a new commit? That costs one Kudu leg instead of the whole matrix, and the sibling leg consistently finishes in about 27 minutes. 2. In parallel I pushed the identical commit `2dd45e26f8613e0ee682d48aed9e963fc1d004f7` to my fork so the full `Build` workflow runs the same SHA once more under a fork where I can re-run individual jobs: https://github.com/DanielLeens/seatunnel/actions/runs/34377408910. Same commit, same tree, so a green result there is a direct check of this head. I will link the outcome here once it completes. @SEZ9 the above is the evidence behind "cancelled, not failed" on `2dd45e26f`. Making the hang fail fast instead of eating 90 minutes is the #12132 fix (a bounded wait in `executeJob` plus a JobManager thread dump when the job never reaches RUNNING) and belongs in its own PR, not here. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
