DanielLeens opened a new issue, #11667: URL: https://github.com/apache/seatunnel/issues/11667
## Background SeaTunnel already displays the current error on the job page, but only as a single raw string: `job.errorMsg` is rendered as one unstructured `<pre>` block (`seatunnel-engine-ui/src/views/jobs/detail.tsx`), with no per-task attribution and no history. Internally the story is the same: `PhysicalPlan.errorBySubPlan` is an `AtomicReference<String>` that uses `compareAndSet(null, ...)` to capture only the *first* error for a sub-plan and silently discards every error after that (`seatunnel-engine/.../dag/physical/PhysicalPlan.java`). `TaskExecutionState` does carry a `throwableMsg` per status report, but it is transient and never persisted per attempt, and there is currently no restart/attempt counter anywhere in the engine (`restartCount`/`attemptNumber`/`retryCount` do not exist). Closely related: every pipeline and task group already maintains a timestamped state-transition history (CREATED to DEPLOYING to RUNNING to FINISHED/CANCELING/CANCELED/FAILING/FAILED) in a `stateTimestamps` array backed by a shared internal map (`SubPlan.java`, `PhysicalVertex.java`), but only the **job-level** timestamps are ever read back (`JobMaster#getStateTimestamp`); the pipeline- and task-group-level timestamps are recorded and then never exposed anywhere. That is exactly the data a "when did this task start failing, and how does that compare to its previous attempt" view would need, and it already exists in memory today. Flink's Web UI addresses both needs with its "Exceptions" tab: the root exception plus the full exception history across restarts, each entry timestamped and attributed to a task/host. ## Problem to solve Because only the first error per sub-plan is kept and there is no attempt concept, diagnosing an intermittently-failing job means either catching the error in real time before something else overwrites the single stored value, or digging through raw logs across every affected worker by hand. This makes it hard to tell a one-off failure apart from a recurring pattern. ## Proposed scope Add a dedicated exception/failure history view, scoped to a job: - persist more than the first error per sub-plan/task: keep a bounded history of failures (timestamp, failing task/subtask, host/worker, exception type and message) - introduce an attempt/restart counter so history entries can be grouped by attempt, since none exists today - expose the already-tracked-but-unread pipeline/task-group `stateTimestamps` alongside each failure entry, so a failure can be shown against how long that attempt had been running rather than as an isolated event - apply to both running and finished jobs, within Zeta's existing history retention window - link each entry to the corresponding task's log (building on #9050/#11662) ## Why this should be a dedicated feature This requires retaining more than "the current error," which has retention, storage, and API-contract implications: - how many historical failures are kept per job, and for how long after the job finishes - how this interacts with existing finished-job history storage (including the known S3-backed finished-jobs issue, #10039) so it does not regress at scale - exposing the pipeline/task-group `stateTimestamps` is extending an existing internal model to a new read path, not new data collection, but it still needs a defined contract ## STIP requirement before implementation Because this changes what failure data Zeta retains and exposes as a stable contract, **the claimant should submit a STIP design first and get maintainer agreement before starting implementation**. The STIP should clarify: - the exception-history data model, its retention limits for running vs. finished jobs, and how attempts are numbered - how large exception messages/stack traces are truncated or paginated - how this scales for jobs stored in external history backends (e.g. S3), consistent with the constraints already surfaced in #10039 - which fields are stable API/UI contract versus best-effort telemetry ## Acceptance criteria - Users can see the full sequence of failures/restarts for a job, not only the latest one, from both the running and finished job views. - Each entry identifies the failing task/attempt/host and links to its log. - Recurring root causes across attempts are visually distinguishable from one-off failures. - English and Chinese docs are updated. ## Non-goals for the first version - automatic root-cause classification or suggested fixes - indefinite retention of full stack traces for all historical jobs - cross-job failure correlation/analytics ## Related work - Issue #9105 (closed) - exception message formatting - Issue #9050 (closed) - worker log file viewing - Issue #11662 - job log links currently broken - Issue #10039 (closed) - finished-jobs history at scale on S3 - Issue #11351 -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
