CryoThrust commented on issue #12139:
URL: https://github.com/apache/seatunnel/issues/12139#issuecomment-5556476068

   I traced the reported call chain in dev: NotifyTaskRestoreOperation catches 
the restoreState failure and sends CheckpointErrorReportOperation; its 
runInternal calls CheckpointManager.reportCheckpointErrorFromTask synchronously 
on the Hazelcast operation thread. CheckpointCoordinator.handleCoordinatorError 
then calls JobMaster.handleCheckpointError, which synchronously transitions 
SubPlan to CANCELING and can enter the restore/state-process path.
   
   The important invariant seems to be that a remote operation must report the 
error and return promptly; pipeline cancellation, resource release, and 
restore/retry should be serialized on the JobMaster or checkpoint coordinator 
executor. A likely shape is to keep the existing single-terminal guard in 
CheckpointCoordinator, then enqueue the state transition exactly once on its 
executor (or a JobMaster-owned executor), rather than invoking SubPlan state 
changes inline from CheckpointErrorReportOperation. Simply making the operation 
fire-and-forget without preserving that guard could race duplicate error 
reports.
   
   A focused regression should make a source enumerator throw during 
restoreState and assert both: (1) the Hazelcast operation completes without 
waiting for restore/retry work, and (2) after the configured retry budget the 
pipeline reaches FAILED and the job future completes with the original 
privilege error. This would distinguish the operation-thread deadlock from the 
connector exception itself.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to