DanielLeens commented on PR #12130:
URL: https://github.com/apache/seatunnel/pull/12130#issuecomment-5605329874

   Follow-up on the CI side, since a clean `Build` run is the only thing still 
open here. I pulled the job logs rather than going by the run summary.
   
   **What is actually red.** On fork run [`awsomesud347/seatunnel` 
34089819303](https://github.com/awsomesud347/seatunnel/actions/runs/34089819303)
 the only job that never went green across attempts 1-4 is `kudu-connector-it 
(11)`, cancelled at the 90-minute job limit every time. The three jobs that 
failed in attempt 1 (`all-connectors-it-4 (11)`, `paimon-connector-it (8)`, 
`rocketmq-connector-it (8)`) all passed on re-run with no code change, and 
`kudu-connector-it (8)` passed in attempt 1 in 27 minutes.
   
   - attempt 1, [job 
101641543194](https://github.com/awsomesud347/seatunnel/actions/runs/34089819303/job/101641543194):
 `KuduIT#testKuduMultipleReadWithRegex` on the Flink 1.15.3 leg. The 
TaskManager logs `The heartbeat of JobManager ... timed out` at 06:48:06, then 
`Could not resolve JobManager address ...` every 10 s until it exits at 
06:53:21 with `RegistrationTimeoutException`; after that the log is silent for 
71 minutes because the test JVM is still inside `container.executeJob`.
   - attempt 4, [job 
101844792732](https://github.com/awsomesud347/seatunnel/actions/runs/34089819303/job/101844792732):
 `KuduIT#testKuduMultipleRead` on the Flink 1.18.0 leg. The JobMaster logs 
`Deploying Source: Kudu-Source -> MultiTableSink-Sink: Writer (1/1) (attempt 
#0)` at 19:39:48 and never logs again. The TaskManager logs `BlobClient - 
Downloading .../p-... from jobmanager` at 19:39:48, `Slot offering to 
JobManager did not finish in time` at 19:39:58 and `The heartbeat of JobManager 
... timed out` at 19:41:48. From there until the cancel at 20:47 the 
ResourceManager/TaskManager slot-allocation loop just repeats (446 
`DefaultSlotStatusSyncer` and 753 `DefaultJobLeaderService` lines) and `flink 
run` never returns.
   
   That is the signature tracked in #12132: the JobMaster stops answering at 
the first task deployment of the first job on a freshly started 
JobManager/TaskManager pair, on the 1.15.3 and 1.18.0 legs in PR runs (and on 
the 1.16.0 and 1.17.2 legs in the scheduled full matrix), never so far on 
1.13.6 or 1.20.1, and `AbstractTestFlinkContainer.executeJob` has no upper 
bound, so the job burns the whole budget. It is not specific to this branch, 
this commit or JDK 11:
   
   - the same hang hit the **JDK 8** leg of this PR's earlier run ([run 
34001671707 attempt 1, job 
101401900552](https://github.com/awsomesud347/seatunnel/actions/runs/34001671707/job/101401900552):
 Flink 1.18.0, `testKuduWholeDatabaseRead`, deploy at 01:23:18, heartbeat 
timeout at 01:25:18) while the JDK 11 leg passed that time;
   - on 2026-09-09 it took both Kudu legs of #12203 ([DanielLeens/seatunnel run 
34207958769](https://github.com/DanielLeens/seatunnel/actions/runs/34207958769),
 same deploy-then-heartbeat-timeout trace on Flink 1.18.0), apache's own 
scheduled dev run 
[34044345096](https://github.com/apache/seatunnel/actions/runs/34044345096) 
lost both Kudu legs the same way on 2026-09-06 (Flink 1.16.0 and 1.17.2 legs, 
deploy then heartbeat timeout 120 s later), and #12169 and #12165 each lost one 
Kudu leg the same way.
   
   Nothing in this diff can reach that path: the change is confined to 
`PageBaseServlet`, `JobInfoService` and `FinishedJobsServlet` in 
`seatunnel-engine-server`, and none of `connector-kudu`, 
`seatunnel-e2e-common`, `connector-kudu-e2e` or the Flink starters declare a 
dependency on that module; the Flink e2e containers only get the Flink starter 
and the connector jars.
   
   **Two things to move this along:**
   
   1. @awsomesud347 only the fork owner can re-run a job on your fork. Could 
you re-run *just the failed job* on run 34089819303 (Actions, open the run, 
"Re-run failed jobs") rather than pushing a new commit? That costs one Kudu leg 
instead of the whole matrix, and the sibling leg consistently finishes in about 
27 minutes.
   2. In parallel I pushed the identical commit 
`2dd45e26f8613e0ee682d48aed9e963fc1d004f7` to my fork so the full `Build` 
workflow runs the same SHA once more under a fork where I can re-run individual 
jobs: https://github.com/DanielLeens/seatunnel/actions/runs/34377408910. Same 
commit, same tree, so a green result there is a direct check of this head. I 
will link the outcome here once it completes.
   
   @SEZ9 the above is the evidence behind "cancelled, not failed" on 
`2dd45e26f`. Making the hang fail fast instead of eating 90 minutes is the 
#12132 fix (a bounded wait in `executeJob` plus a JobManager thread dump when 
the job never reaches RUNNING) and belongs in its own PR, not here.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to