goutamadwant opened a new issue, #11790:
URL: https://github.com/apache/seatunnel/issues/11790

   ### Search before asking
   
   - [x] I searched the existing feature and design issues. The operator-facing 
timeline is tracked by #11355. The related runtime ordering and recovery 
contract is tracked separately by #11402.
   
   ### Description
   
   ## Summary
   
   This STIP proposes a bounded schema evolution timeline and decision log for 
SeaTunnel CDC jobs.
   
   The first version is a runtime observability feature. It is not a durable 
compliance audit log and does not redefine schema-change correctness. It 
records the normalized decisions and outcomes produced by the existing CDC, 
transform, routing, engine, and sink paths so operators can understand what 
happened to a schema change.
   
   ## Current behavior
   
   SeaTunnel carries `SchemaChangeEvent` objects through the source collector, 
transform chain, and sink lifecycle. Production sink writers can apply schema 
changes, and multi-table writers coordinate schema-change barriers across their 
workers.
   
   The current paths log parts of this lifecycle independently. They do not 
provide one job-scoped read model that answers:
   
   - which schema event was observed
   - how its table identity changed through routing or transforms
   - why it was forwarded, filtered, ignored, rejected, or applied
   - which sink targets were involved
   - whether every target completed successfully
   - whether the information is still available after the job finishes
   
   ## Relationship to the runtime correctness contract
   
   Issue #11402 owns the ordering, schema epoch, checkpoint, replay, restore, 
and unsupported-path correctness contract.
   
   This proposal consumes that contract. It does not introduce a second schema 
state machine or make observability storage part of checkpoint correctness.
   
   If a schema mutation succeeds but the runtime later restores or replays the 
event, the behavior must follow #11402. The timeline reports the observed 
attempt and outcome. A failure to write timeline data must never change the 
schema mutation or restore result.
   
   ## Goals
   
   1. Expose a bounded timeline of recent schema evolution activity for running 
and finished jobs.
   2. Correlate framework-owned decisions for one normalized schema event.
   3. Represent source table, resolved target table, event type, policy 
decision, sink capability, and final outcome.
   4. Represent multi-table fan-out without hiding partial application.
   5. Distinguish stable product fields from optional connector debug details.
   6. Keep raw DDL and sensitive payloads out of the default public contract.
   7. Preserve existing job, connector, and schema-change behavior when 
observability is unavailable.
   
   ## Non-goals
   
   The first version will not provide:
   
   - indefinite event retention
   - a compliance or forensic audit guarantee
   - full raw DDL archival
   - automatic repair or schema recommendations
   - identical connector-specific detail for every source and sink
   - a new ordering, replay, or checkpoint protocol
   - cross-job schema analytics
   
   ## Design principles
   
   ### Framework-owned facts are stable
   
   The stable model should contain only facts SeaTunnel can define consistently:
   
   - event correlation ID
   - job ID and sequence
   - source and target table identifiers
   - normalized schema event type
   - decision stage
   - configured policy when available
   - sink capability result
   - target outcome and reason code
   - timestamps and truncation flags
   
   Connector-specific offsets, native DDL fragments, vendor error fields, and 
payload excerpts are optional debug details. Clients must not depend on them 
being present across connectors or releases.
   
   ### Observability is best effort
   
   Timeline recording must not block, fail, retry, or acknowledge schema 
changes. Recorder and storage failures are logged and counted, while the 
original schema-change path continues with its existing result.
   
   ### No raw DDL by default
   
   `TableEvent.statement` and connector-native payloads can contain names, 
literals, or other sensitive data. V1 does not expose raw DDL through the REST 
response. It records the normalized event type and bounded structural metadata. 
A future opt-in raw payload design would require separate security and 
retention review.
   
   ## Event model
   
   One logical normalized schema event produces one `SchemaEvolutionRecord`.
   
   Proposed stable fields:
   
   - `schemaChangeId`: stable correlation ID assigned when SeaTunnel first 
normalizes the event
   - `sequence`: job-scoped monotonically increasing sequence assigned by the 
history owner
   - `jobId`
   - `eventCreatedAt`: source event timestamp when available
   - `observedAt`
   - `completedAt`: optional until terminal
   - `sourceTable`
   - `eventType`
   - `state`: `IN_PROGRESS` or `TERMINAL`
   - `outcome`: optional until terminal
   - `policy`: configured behavior when available, otherwise `LEGACY_DEFAULT` 
or `NOT_EVALUATED`
   - `decisions`: ordered framework decision entries
   - `targets`: zero or more sink target decisions
   - `detailsTruncated`
   
   The exact carrier for `schemaChangeId` must be agreed before implementation. 
It may be additive event metadata or an internal envelope, but it must survive 
task transport and replay. It must not rely only on object identity or a 
timestamp-derived key.
   
   ## Decision stages
   
   The first version records these normalized stages when they occur:
   
   1. `OBSERVED`: the source emitted a schema change into the SeaTunnel runtime.
   2. `NORMALIZED`: the connector event became a recognized SeaTunnel 
`SchemaChangeEvent`.
   3. `TRANSFORMED`: a transform changed the event or target schema.
   4. `FILTERED`: a source rule or transform intentionally removed the event.
   5. `TARGET_RESOLVED`: source-to-target table resolution completed.
   6. `POLICY_EVALUATED`: configured schema behavior was evaluated when that 
feature is available.
   7. `CAPABILITY_EVALUATED`: sink support was determined as `SUPPORTED`, 
`UNSUPPORTED`, or `UNKNOWN`.
   8. `APPLY_STARTED`: sink-side application began for a target.
   9. `COMPLETED`: the target reached a terminal outcome.
   
   Not every connector or path will emit every stage. Missing stages are 
represented as unavailable rather than inferred.
   
   ## Outcomes and reason codes
   
   Proposed record outcomes:
   
   - `APPLIED`
   - `IGNORED`
   - `FILTERED`
   - `FAILED`
   - `PARTIALLY_APPLIED`
   
   The outcome is separate from a stable reason code. Initial reason codes 
should cover:
   
   - `POLICY_STRICT`
   - `POLICY_IGNORE`
   - `SOURCE_EVENT_FILTER`
   - `TRANSFORM_EVENT_FILTER`
   - `TARGET_RESOLUTION_FAILED`
   - `SINK_UNSUPPORTED`
   - `SINK_APPLY_FAILED`
   - `APPLY_SUCCEEDED`
   - `UNKNOWN`
   
   New reason codes may be added without changing existing meanings. Clients 
must handle unknown values.
   
   ## Multi-table and routed events
   
   A single source event can resolve to one or more sink targets. The parent 
record retains the source event identity and overall outcome. Each target entry 
contains:
   
   - resolved target table
   - sink or sub-writer identity when reliably available
   - capability result
   - apply start and completion timestamps
   - target outcome
   - reason code
   - bounded error summary when failed
   
   The parent is `APPLIED` only when every required target is applied. It is 
`PARTIALLY_APPLIED` when at least one externally visible target mutation 
succeeded and another target failed or remained unknown.
   
   The target list must be bounded. If a single event exceeds the configured or 
fixed V1 bound, the response records the total target count and marks the 
returned target details as truncated.
   
   ## Deduplication and replay
   
   The timeline is diagnostic and must not drive correctness decisions.
   
   Updates are merged idempotently by `schemaChangeId + stage + target 
identity`. Repeated delivery of the same stage does not create another decision 
entry.
   
   A replay of the same serialized schema event keeps the same `schemaChangeId` 
and can append a new attempt marker or replay decision. A newly observed source 
event receives a new ID even when its event type and table match an earlier 
event.
   
   The recorder must not guess that two events are the same from table, type, 
statement, or timestamp alone.
   
   ## Storage and retention
   
   The history owner keeps a bounded job-scoped collection in HA-backed engine 
state.
   
   Proposed V1 behavior:
   
   - retain at most 500 logical schema event records per job
   - evict the oldest terminal record first
   - do not evict an in-progress record unless the job itself is removed
   - cap target decisions at 100 per logical event
   - cap each optional error summary at 4 KiB of valid UTF-8
   - retain records while the job is active
   - write one bounded snapshot when the job becomes terminal
   - expire the snapshot with the existing finished-job history lifecycle
   
   These limits are initial safety bounds, not a long-term audit guarantee. The 
exact values should be confirmed using operational evidence during review.
   
   Storage and merge failures are best effort. They must not alter the job 
state, schema event, sink apply result, or checkpoint behavior.
   
   ## REST contract
   
   Proposed endpoint:
   
   `GET /job-info/{jobId}/schema-evolution`
   
   Query parameters:
   
   - `limit`: number of logical events, default 100 and capped at the retained 
maximum
   - `beforeSequence`: cursor for older records
   - `table`: optional exact source or target table filter
   - `outcome`: optional normalized outcome filter
   
   Behavior:
   
   - return records in descending sequence order
   - return the same response model for running and finished jobs
   - return an empty collection for a known job with no records
   - use existing job-not-found behavior for unknown or expired jobs
   - expose truncation and unavailable fields explicitly
   - do not return raw DDL or arbitrary connector payloads
   
   The existing job detail and error fields remain unchanged.
   
   ## Web UI
   
   The Web UI follows after the backend model is stable.
   
   The job detail page should:
   
   - show recent schema events in time order
   - group per-target decisions under the parent event
   - show source and resolved target tables
   - show event type, policy, capability, final outcome, and reason
   - distinguish ignored, filtered, unsupported, failed, and partially applied 
states
   - mark unavailable or truncated details
   - avoid presenting the timeline as a durable audit record
   
   ## Compatibility
   
   The proposal is additive:
   
   - existing jobs require no configuration changes
   - schema-change ordering and sink behavior remain unchanged
   - existing REST fields remain unchanged
   - recorder failures do not affect job correctness
   - finished-job expiration follows the existing lifecycle
   - connectors may initially provide only framework-owned fields
   
   If adding a correlation ID changes serialized event state, compatibility 
with existing checkpoints and savepoints must be tested. Incompatible state 
must not be introduced silently.
   
   ## Implementation slices
   
   1. Add the design documentation and agree on the stable model.
   2. Add internal record, recorder, bounded active-job storage, and 
finished-job snapshot.
   3. Add correlation and source/transform decision capture.
   4. Add routing, policy, capability, and per-target sink outcome capture.
   5. Add REST routing and compatibility tests.
   6. Add the Web UI after the REST model is stable.
   7. Add E2E recovery and multi-target validation plus English and Chinese 
operational docs.
   
   ## Validation plan
   
   Unit and integration coverage should include:
   
   - successful source normalization and sink apply
   - source-filtered and transform-filtered events
   - source-to-target table rename or routing
   - supported and unsupported sink capability decisions
   - sink apply failure
   - multi-target success and partial application
   - repeated stage delivery without duplicate decisions
   - replay preserving the correlation ID
   - active master failover preserving bounded records
   - oldest-terminal eviction and target truncation
   - UTF-8-safe detail truncation
   - recorder/store failure not changing schema behavior
   - identical REST shape for running and finished jobs
   - backward compatibility for existing job-info endpoints
   
   At least one E2E path should use MySQL CDC to a schema-evolution-capable 
JDBC sink. A negative path should use an unsupported sink or unsupported event 
type.
   
   ## Acceptance criteria
   
   1. Operators can inspect recent schema evolution activity for running and 
finished jobs.
   2. One logical event is correlated across the framework-owned decision 
stages that are available.
   3. Source table, resolved targets, event type, capability, and final outcome 
are visible.
   4. Filtered, ignored, unsupported, failed, and partially applied outcomes 
are distinguishable.
   5. Multi-target decisions do not collapse into a misleading single success 
state.
   6. Records, target lists, and optional error details are bounded and expose 
truncation.
   7. Raw DDL is not exposed by default.
   8. Timeline failures do not affect schema-change or job correctness.
   9. Existing job-info and schema evolution behavior remain backward 
compatible.
   10. Correlation and retained records survive the supported task transport 
and active-master failover paths.
   11. English and Chinese documentation describe the model, retention, privacy 
boundary, and connector limitations.
   
   ## Usage Scenario
   
   An operator running a CDC job sees that a target table no longer matches the 
source. Instead of reconstructing the path from distributed logs, the operator 
opens the job's schema evolution timeline and can answer:
   
   - when SeaTunnel observed the schema change
   - which normalized event type was produced
   - whether a source rule or transform filtered it
   - how the source table resolved to one or more target tables
   - which behavior policy was evaluated
   - whether each target supported and applied the change
   - whether the event was ignored, failed, or partially applied
   - whether any optional details were unavailable or truncated
   
   ## Related work
   
   - Feature tracking issue: #11355
   - Schema ordering and recovery contract: #11402
   - Schema behavior policies: #11043
   - Schema event type filtering: #11044
   - Related implementation work: #11015, #11025, #11162, #11328
   
   ## Are you willing to submit a PR?
   
   - [x] Yes I am willing to submit a PR!
   
   ## Code of Conduct
   
   - [x] I agree to follow this project's [Code of 
Conduct](https://www.apache.org/foundation/policies/conduct)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to