goutamadwant opened a new issue, #11735:
URL: https://github.com/apache/seatunnel/issues/11735

   ### Search before asking
   
   - [x] I had searched in the 
[feature](https://github.com/apache/seatunnel/issues?q=is%3Aissue+label%3A%22Feature%22)
 and found no similar feature requirement.
   
   
   ### Description
   
   ## Summary
   
   This STIP proposes a bounded and structured task failure history for Zeta 
Engine.
   
   Today, SeaTunnel retains and exposes a single formatted error message for a 
job. This is insufficient when a pipeline is restored multiple times or when 
different task groups fail during the same job.
   
   The proposal introduces a job-scoped failure history that:
   
   - groups failures by pipeline execution attempt
   - identifies the failing pipeline and task group
   - preserves worker and task information when available
   - records timestamps, exception type, message, and stack trace
   - remains available for running and finished jobs
   - survives active master changes
   - expires through the existing finished-job history policy
   - preserves the existing `errorMessage` contract for compatibility
   
   ## Current behavior
   
   `PhysicalPlan` currently retains only the first error reported by a sub-plan.
   
   `TaskExecutionState` transports a formatted throwable message, but the 
engine does not preserve structured failure records across retries or restores.
   
   `JobHistoryService` retains the final job error for finished jobs, but not 
the sequence of failures that led to the terminal state.
   
   Pipeline and task-group state timestamps already exist internally, but they 
are not exposed as part of a failure-history read path.
   
   ## Goals
   
   1. Retain multiple failures for one job using a bounded history.
   2. Group failures by pipeline execution attempt.
   3. Preserve task-group and worker attribution where available.
   4. Support running and finished jobs through the same read contract.
   5. Preserve history across active master changes.
   6. Define deterministic retention, ordering, and deduplication behavior.
   7. Keep existing job-detail clients backward compatible.
   8. Provide a stable backend contract before implementing the Web UI.
   
   ## Non-goals
   
   The first version will not provide:
   
   - automatic root-cause classification
   - cross-job failure analytics
   - indefinite stack-trace retention
   - distributed log aggregation
   - task-level attribution when only task-group information is available
   - changes to checkpoint or savepoint payloads
   
   ## Attempt model
   
   An attempt belongs to a pipeline rather than an individual task.
   
   - Initial pipeline execution uses attempt `0`.
   - The attempt number increments before pipeline restore is scheduled.
   - Failures from the same pipeline execution carry the same attempt number.
   - Attempt state must survive active master changes.
   - A restored pipeline receives a new attempt even when an individual task 
index is reused.
   
   This follows the existing pipeline restore boundary without introducing new 
task-level retry semantics.
   
   ## Failure record
   
   A failure record contains:
   
   - sequence
   - timestamp
   - job ID
   - pipeline ID
   - pipeline attempt
   - attempt start time
   - task-group ID
   - task ID when available
   - task name when available
   - worker address when available
   - exception type when transported structurally
   - concise exception message
   - stack trace
   
   The required fields are:
   
   - sequence
   - timestamp
   - job ID
   - pipeline ID
   - attempt
   - task-group ID
   
   Task ID, task name, worker, exception type, message, and stack trace remain 
optional because not every existing failure path provides them.
   
   The implementation must not derive exception type or task identity by 
parsing formatted stack traces.
   
   ## Capture and deduplication
   
   `TaskExecutionState` is the existing task-to-master failure transport 
boundary. The implementation should extend this transport with structured 
failure fields.
   
   The JobMaster captures the failure before task resources are released.
   
   Repeated delivery of the same terminal state must not create duplicate 
records. The first implementation will deduplicate using:
   
   `pipelineId + attempt + taskGroupId`
   
   A task group that fails again after restore has a new attempt and therefore 
produces a separate record.
   
   Failure-history persistence is diagnostic and best effort. A history-store 
failure must not block the original failure or restore processing.
   
   ## Storage and retention
   
   Failure history should use an HA-backed engine state-store abstraction keyed 
by job ID.
   
   The concrete Hazelcast map or an external history backend remains an 
implementation detail.
   
   Initial retention behavior:
   
   - retain at most 100 records per job
   - evict the oldest record when the limit is exceeded
   - do not expire records while a job is active
   - apply `history-job-expire-minutes` after the job reaches a terminal state
   - remove associated history when the finished-job record expires
   
   The initial maximum is a fixed implementation bound. A configurable limit 
can be considered later based on operational evidence.
   
   ## Large exception handling
   
   Exception payloads must be bounded independently from the number of records.
   
   The implementation should define:
   
   - a maximum message size
   - a maximum stored stack-trace size
   - explicit truncation metadata
   - truncation at a valid character boundary
   - preservation of the beginning and end of the stack trace when practical
   
   The REST response should expose whether a field was truncated.
   
   The first version will not paginate an individual stack trace. Future 
versions may add a separate detail endpoint without changing the failure 
summary contract.
   
   The exact byte limits should be finalized with maintainer input before 
implementation.
   
   ## External history backends
   
   The public contract must not depend directly on Hazelcast `IMap`.
   
   Running-job history may use the active engine state store. When a job 
finishes, its bounded failure history must follow the same lifecycle as the 
finished-job record.
   
   External history implementations such as S3 must be able to persist the same 
bounded representation without requiring an unbounded object or one object per 
exception.
   
   The first implementation should write one bounded failure-history snapshot 
with the finished-job history record or through the same history-store 
abstraction.
   
   Backend write failures must follow existing finished-job history error 
handling and must not change the job terminal state.
   
   ## REST contract
   
   Proposed endpoint:
   
   `GET /job-info/{jobId}/failures?limit=100`
   
   Behavior:
   
   - return records in descending sequence order
   - default `limit` to 100
   - reject non-positive limits
   - cap the requested limit at the retained maximum
   - return an empty list for a known job with no recorded failures
   - use existing job-not-found behavior for unknown or expired jobs
   - use the same response model for running and finished jobs
   
   The existing `errorMessage` field remains unchanged.
   
   ## Web UI
   
   The Web UI will be implemented separately after the backend contract is 
stable.
   
   The Exception view should:
   
   - group failures by pipeline attempt
   - display timestamp, pipeline, task group, worker, exception type, and 
message
   - collapse stack traces by default
   - indicate truncated fields
   - link to task logs when a reliable log link is available
   - display missing optional fields as unavailable
   - avoid claiming task-level precision for task-group-level failures
   
   ## Compatibility
   
   This proposal is additive:
   
   - existing jobs require no configuration changes
   - existing REST fields remain unchanged
   - the existing final `errorMessage` remains available
   - checkpoint and savepoint formats remain unchanged
   - old failure paths may populate only the fields available to them
   
   ## Implementation slices
   
   1. Add the internal failure record and HA-backed storage abstraction.
   2. Add pipeline-attempt persistence and structured task failure transport.
   3. Add capture, deduplication, retention, truncation, restore, and failover 
tests.
   4. Add finished-job history backend integration.
   5. Add the REST endpoint and API tests.
   6. Add the Web UI history view separately.
   
   ## Acceptance criteria
   
   1. A first-attempt failure is recorded with attempt `0`.
   2. A failure after restore is recorded under the incremented attempt.
   3. Duplicate terminal-state delivery does not create duplicate records.
   4. Different task-group failures remain separate.
   5. Active master failover preserves records and attempt numbering.
   6. Running and finished jobs expose the same response model.
   7. More than 100 records evicts the oldest records deterministically.
   8. Large exception fields are bounded and marked as truncated.
   9. External history backends can retain the bounded representation.
   10. His
   ```
   
   ### Usage Scenario
   
   This feature is intended for operators diagnosing Zeta jobs that fail or 
restore more than once.
   
   It should help answer:
   
   - Which pipeline attempt failed?
   - Which task group and worker reported the failure?
   - Is the same root cause recurring across restores?
   - How long had the attempt been running before it failed?
   - Which failures occurred before the final job error?
   - Is the full stack trace available or was it truncated?
   - Can the same history be inspected after the job finishes?
   
   ### Related issues
   
   Feature umbrella: #11667
   
   Design pull request: #11734
   
   Related work:
   
   - #9105
   - #9050
   - #10039
   - #11351
   - #11662
   
   ### Are you willing to submit a PR?
   
   - [x] Yes I am willing to submit a PR!
   
   ### Code of Conduct
   
   - [x] I agree to follow this project's [Code of 
Conduct](https://www.apache.org/foundation/policies/conduct)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to