DanielLeens opened a new issue, #12139: URL: https://github.com/apache/seatunnel/issues/12139
### Search before asking - [x] I had searched in the [issues](https://github.com/apache/seatunnel/issues?q=is%3Aissue) and found no similar issues. ### What happened `PaimonIT#privilegeEnabledPaimonSourceUnAuthorized` (Zeta leg) submits `paimon_to_paimon_privilege1.conf` and expects the job to **fail** (exit code 1) because the user has no SELECT privilege. Twice today the job neither failed nor finished: Zeta entered its pipeline-restore loop, the restore re-threw the same `NoPrivilegeException` from inside `SourceSplitEnumeratorTask.restoreState`, and after that the JobMaster went silent for the rest of the job's budget. `seatunnel.sh` never returned, the test blocked in `container.executeJob`, and GitHub cancelled the job at the 180-minute limit. - PR #11077, fork run `hesam-oxe/seatunnel` 33988030045, job `paimon-connector-it (11)` (job 101365445865), 21:21 -> 00:21 UTC; job failed at 21:47:34, last restore step 21:47:37, silence until the cancel. - PR #11503, fork run `nielifeng/seatunnel` 33994603041, job `paimon-connector-it (8)` (job 101383153941), 23:42 -> 02:42 UTC; job failed at 00:10:46, last restore step 00:10:49, silence until the cancel. Neither PR touches connector-paimon or the engine restore path. ### Timeline (run 33994603041; the other run is identical) ``` 00:10:46 CheckpointCoordinator - report error from task 00:10:46 SeaTunnelException: org.apache.paimon.privilege.NoPrivilegeException: User paimon doesn't have privilege SELECT on table default.st_test_p 00:10:46 CheckpointCoordinator - checkpoint_state_<job>_1 has already been cleaned, skip persisting transition to FAILED 00:10:46 SubPlan - Task TaskGroupLocation{..., taskGroupId=3} Failed in Job paimon_to_paimon_privilege1.conf, Pipeline: [(1/1)] 00:10:46 CheckpointCoordinator - start clean pending checkpoint cause CheckpointCoordinator inside have error. 00:10:46 SubPlan - Restore time 1, pipeline Job paimon_to_paimon_privilege1.conf, Pipeline: [(1/1)] 00:10:46 SubPlan - Reset pipeline ... state to CREATED 00:10:46 CheckpointCoordinator - received restore CheckpointCoordinator with alreadyStarted: false 00:10:46 SubPlan - Wait 3s and then restore the pipeline ... 00:10:49 (restore) NotifyTaskRestoreOperation -> SourceSplitEnumeratorTask.restoreState(SourceSplitEnumeratorTask.java:201) -> PaimonSource.createEnumerator(PaimonSource.java:155) -> AbstractSplitEnumerator.<init>(AbstractSplitEnumerator.java:120/124) -> ReadBuilderImpl.newScan -> PrivilegedFileStoreTable.newScan -> NoPrivilegeException (same as above) -> CheckpointErrorReportOperation.runInternal(:48) -> CheckpointManager.reportCheckpointErrorFromTask(:239) -> CheckpointCoordinator.reportCheckpointErrorFromTask(:519) -> handleCoordinatorError(:352) -> CheckpointManager.handleCheckpointError(:231) -> JobMaster.handleCheckpointError(:773/776) -> SubPlan.handleCheckpointError(SubPlan.java:644) 00:11:00 com.hazelcast...SlowOperationDetector - Slow operation detected: CheckpointErrorReportOperation (stack ending in SubPlan.handleCheckpointError(SubPlan.java:644)) ... no further job/pipeline state transition until the cancel at 02:42 ``` In the first sighting the loop got one step further (`Restore time 2` at 21:47:37) before the same stall. ### Root cause (from `dev` source) - `PaimonSource#createEnumerator` -> `AbstractSplitEnumerator` constructor calls `newScan()` on every table eagerly, so a privilege failure is thrown from **`createEnumerator` itself**, i.e. from `SourceSplitEnumeratorTask.restoreState` during the pipeline restore rather than from `run()`. - That exception is reported back through `CheckpointErrorReportOperation`, which runs on a Hazelcast operation thread and synchronously drives `CheckpointCoordinator.handleCoordinatorError` -> `JobMaster.handleCheckpointError` -> `SubPlan.handleCheckpointError`. The `SlowOperationDetector` warning shows that call never returns. - `SubPlan.prepareRestorePipeline()` (SubPlan.java:489) holds `restoreLock` while it resets the pipeline and sleeps `job.retry.interval.seconds`, and the pipeline end callback path (`checkNeedRestore` / `prepareRestorePipeline` / `restorePipeline`, SubPlan.java:739-742) is what should drive the next restore attempt (`job.retry.times` defaults to 3). With the error handler blocked on the operation thread, the pipeline never reaches an end state again: no further `Restore time N`, no `FAILED`, no job completion. The retry budget is never exhausted because the retry loop itself is stuck. ### What you expected to happen A job whose source cannot even be created on restore should fail after `job.retry.times` attempts (here: exit code 1 within seconds), not hang indefinitely. The engine must not run pipeline error handling synchronously on a Hazelcast operation thread in a way that can block the restore loop, and a failing `createEnumerator` during `restoreState` must be treated like any other task failure. ### SeaTunnel Version dev (`af0a647d`), Zeta. ### Engine Zeta ### Additional context Each occurrence costs a full 180-minute CI job. Filed while triaging CI for #11077 and #11503. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
