DanielLeens opened a new issue, #12139:
URL: https://github.com/apache/seatunnel/issues/12139

   ### Search before asking
   
   - [x] I had searched in the 
[issues](https://github.com/apache/seatunnel/issues?q=is%3Aissue) and found no 
similar issues.
   
   ### What happened
   
   `PaimonIT#privilegeEnabledPaimonSourceUnAuthorized` (Zeta leg) submits 
`paimon_to_paimon_privilege1.conf` and expects the job to **fail** (exit code 
1) because the user has no SELECT privilege. Twice today the job neither failed 
nor finished: Zeta entered its pipeline-restore loop, the restore re-threw the 
same `NoPrivilegeException` from inside 
`SourceSplitEnumeratorTask.restoreState`, and after that the JobMaster went 
silent for the rest of the job's budget. `seatunnel.sh` never returned, the 
test blocked in `container.executeJob`, and GitHub cancelled the job at the 
180-minute limit.
   
   - PR #11077, fork run `hesam-oxe/seatunnel` 33988030045, job 
`paimon-connector-it (11)` (job 101365445865), 21:21 -> 00:21 UTC; job failed 
at 21:47:34, last restore step 21:47:37, silence until the cancel.
   - PR #11503, fork run `nielifeng/seatunnel` 33994603041, job 
`paimon-connector-it (8)` (job 101383153941), 23:42 -> 02:42 UTC; job failed at 
00:10:46, last restore step 00:10:49, silence until the cancel.
   
   Neither PR touches connector-paimon or the engine restore path.
   
   ### Timeline (run 33994603041; the other run is identical)
   
   ```
   00:10:46 CheckpointCoordinator - report error from task
   00:10:46 SeaTunnelException: 
org.apache.paimon.privilege.NoPrivilegeException: User paimon doesn't have 
privilege SELECT on table default.st_test_p
   00:10:46 CheckpointCoordinator - checkpoint_state_<job>_1 has already been 
cleaned, skip persisting transition to FAILED
   00:10:46 SubPlan - Task TaskGroupLocation{..., taskGroupId=3} Failed in Job 
paimon_to_paimon_privilege1.conf, Pipeline: [(1/1)]
   00:10:46 CheckpointCoordinator - start clean pending checkpoint cause 
CheckpointCoordinator inside have error.
   00:10:46 SubPlan - Restore time 1, pipeline Job 
paimon_to_paimon_privilege1.conf, Pipeline: [(1/1)]
   00:10:46 SubPlan - Reset pipeline ... state to CREATED
   00:10:46 CheckpointCoordinator - received restore CheckpointCoordinator with 
alreadyStarted: false
   00:10:46 SubPlan - Wait 3s and then restore the pipeline ...
   00:10:49 (restore) NotifyTaskRestoreOperation -> 
SourceSplitEnumeratorTask.restoreState(SourceSplitEnumeratorTask.java:201)
                     -> PaimonSource.createEnumerator(PaimonSource.java:155)
                     -> 
AbstractSplitEnumerator.<init>(AbstractSplitEnumerator.java:120/124) -> 
ReadBuilderImpl.newScan -> PrivilegedFileStoreTable.newScan
                     -> NoPrivilegeException (same as above)
                     -> CheckpointErrorReportOperation.runInternal(:48) -> 
CheckpointManager.reportCheckpointErrorFromTask(:239)
                     -> 
CheckpointCoordinator.reportCheckpointErrorFromTask(:519) -> 
handleCoordinatorError(:352)
                     -> CheckpointManager.handleCheckpointError(:231) -> 
JobMaster.handleCheckpointError(:773/776) -> 
SubPlan.handleCheckpointError(SubPlan.java:644)
   00:11:00 com.hazelcast...SlowOperationDetector - Slow operation detected: 
CheckpointErrorReportOperation (stack ending in 
SubPlan.handleCheckpointError(SubPlan.java:644))
      ... no further job/pipeline state transition until the cancel at 02:42
   ```
   
   In the first sighting the loop got one step further (`Restore time 2` at 
21:47:37) before the same stall.
   
   ### Root cause (from `dev` source)
   
   - `PaimonSource#createEnumerator` -> `AbstractSplitEnumerator` constructor 
calls `newScan()` on every table eagerly, so a privilege failure is thrown from 
**`createEnumerator` itself**, i.e. from 
`SourceSplitEnumeratorTask.restoreState` during the pipeline restore rather 
than from `run()`.
   - That exception is reported back through `CheckpointErrorReportOperation`, 
which runs on a Hazelcast operation thread and synchronously drives 
`CheckpointCoordinator.handleCoordinatorError` -> 
`JobMaster.handleCheckpointError` -> `SubPlan.handleCheckpointError`. The 
`SlowOperationDetector` warning shows that call never returns.
   - `SubPlan.prepareRestorePipeline()` (SubPlan.java:489) holds `restoreLock` 
while it resets the pipeline and sleeps `job.retry.interval.seconds`, and the 
pipeline end callback path (`checkNeedRestore` / `prepareRestorePipeline` / 
`restorePipeline`, SubPlan.java:739-742) is what should drive the next restore 
attempt (`job.retry.times` defaults to 3). With the error handler blocked on 
the operation thread, the pipeline never reaches an end state again: no further 
`Restore time N`, no `FAILED`, no job completion. The retry budget is never 
exhausted because the retry loop itself is stuck.
   
   ### What you expected to happen
   
   A job whose source cannot even be created on restore should fail after 
`job.retry.times` attempts (here: exit code 1 within seconds), not hang 
indefinitely. The engine must not run pipeline error handling synchronously on 
a Hazelcast operation thread in a way that can block the restore loop, and a 
failing `createEnumerator` during `restoreState` must be treated like any other 
task failure.
   
   ### SeaTunnel Version
   
   dev (`af0a647d`), Zeta.
   
   ### Engine
   
   Zeta
   
   ### Additional context
   
   Each occurrence costs a full 180-minute CI job. Filed while triaging CI for 
#11077 and #11503.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to