DanielLeens opened a new pull request, #12096: URL: https://github.com/apache/seatunnel/pull/12096
### Purpose of this pull request Fixes #12094. `CheckpointCoordinator` previously rebuilt `completedCheckpointIds` as an empty queue after an active-master failover. Checkpoints written by earlier coordinator instances were therefore invisible to retention cleanup and could accumulate beyond `checkpoint.storage.max-retained`. This patch: - rebuilds the durable checkpoint retention queue for the current job and pipeline when a running coordinator is recreated; - sorts storage enumeration results by checkpoint ID before selecting the oldest files; - enforces the configured maximum after recovery and after every subsequent durable checkpoint; - keeps explicit checkpoint restores to a different destination job isolated from the source job's retention state; - makes LocalFile and HDFS batch deletion failures observable while keeping retention cleanup best-effort at the coordinator layer, so cleanup failures do not fail checkpoint completion or active-master recovery and remain eligible for retry. ### Does this PR introduce _any_ user-facing change? Yes. Previously, a running Zeta job could retain more checkpoint files than `checkpoint.storage.max-retained` after one or more active-master failovers. After this change, the coordinator restores the existing checkpoint IDs and removes the oldest files until the configured bound is satisfied. A transient cleanup failure is logged and retried after the next completed checkpoint without interrupting the job. ### How was this patch tested? Added regression tests covering: - unordered checkpoint enumeration and exact cleanup during active-master coordinator recreation; - exact retention on the next completed checkpoint; - isolation when restoring from a different source job; - cleanup failure not interrupting coordinator recreation or checkpoint completion, with retry on the next checkpoint; - LocalFile delete exceptions and HDFS `FileSystem.delete` returning `false`. Formatting verification completed successfully: ```shell ./mvnw -nsu -pl seatunnel-engine/seatunnel-engine-server,seatunnel-engine/seatunnel-engine-storage/checkpoint-storage-plugins/checkpoint-storage-local-file,seatunnel-engine/seatunnel-engine-storage/checkpoint-storage-plugins/checkpoint-storage-hdfs -am spotless:apply ./mvnw -nsu -pl seatunnel-engine/seatunnel-engine-server,seatunnel-engine/seatunnel-engine-storage/checkpoint-storage-plugins/checkpoint-storage-local-file,seatunnel-engine/seatunnel-engine-storage/checkpoint-storage-plugins/checkpoint-storage-hdfs -am spotless:check ``` Functional tests are delegated to GitHub CI and were not executed locally. ### Check list * [x] No new Jar binary package is added. * [x] Documentation changes are not required because no configuration or public API changes. * [x] `incompatible-changes.md` changes are not required because the patch restores the documented retention behavior. * [x] This PR does not change connector code. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
