DanielLeens commented on PR #12284: URL: https://github.com/apache/seatunnel/pull/12284#issuecomment-5662049411
Follow-up to the second review round and to the fork's Build run for `e7bccdfb0b2` (run 34740882765). Commit `ea03511df99` addresses everything that is this PR's to fix. **Blocker 1, staged hand-off behind a column-appending wrapper: fixed.** The hand-off no longer requires positional equality between the hint-derived schema and the upstream produced table. `SQLLineageSchema.inOrderOf` adopts the upstream layout when it holds exactly the columns the hints produce (same names and definitions), keeping identities, creating hints and repositioning flags; anything else still fails fast with `TRANSFORM_COMMON-09`. For `Metadata -> Sql (select *)` the translator now emits `ADD age AFTER weight` (the last base column), which replays onto the pre-event output, so `ProductionPipelineSchemaChangeTest` passes by construction. `TransformChainLiveAlterTest` drives the transforms without the engine hand-off and keeps its tail positions. New unit tests: `testStagedHandOffAdoptsLayoutOfColumnAppendingWrapper` (Metadata upstream, rows through) and `testStagedHandOffRejectsLayoutWithOtherColumns`. The error message names the event class and type when the event carr ies no statement instead of printing `'null'`. Docs describe the adopted layout. **Blocker 2, `TestFilterRowKindIT.testFilterRowKindMultiTable`: not caused by this PR, but checked rather than dismissed.** Unmatched tables of that job do go through `IdentityMapTransform`, which this PR touches. In the same run the job passed on Zeta, Spark 2.4.6, Spark 3.3.0, Flink 1.13.6 and Flink 1.18.0 and failed only on the Flink 1.15.3 container, where the log shows the Assert sink failing `MIN_ROW 100` with 25 and then 91 rows, the Flink job restarting twice, and then `MAX_ROW 100` failing with 150 rows: counts accumulated across restarts, the `AssertSinkWriter` static-counter problem that has its own fix in flight. An identity regression would fail on every engine. **This PR's own E2E failures: root cause in `MultiTableSinkWriter`, worked around and documented.** With `multi_table_sink_replica = 2` the writer dispatches every schema change to both sub-writers of the same physical table. The second sub-writer replays `CHANGE reuse_col reuse_col_old, ADD reuse_col` after the first one already re-added `reuse_col`, so MySQL answers `Duplicate column name 'reuse_col_old'` and the job restarts forever; `DROP name, ADD name` is executed twice, and the projection test's query happened to run between the two statements, which aborted the test instead of retrying. The same events reach the sink without any transform, so this is independent of the translator; the physical DDL should be applied once per physical table and siblings refreshed at runtime only, the way `RestoreTableSchemaEvent` already is. Recorded as a follow-up in the description. The two SQL jobs now run with one sub-writer per table (the pre-existing jobs keep replica 2), and polls tha t hit the moment between the two statements of one composite are retried. **Remaining failures in that run, none in this PR's scope:** Iceberg S3, S3 file, Paimon S3 and Databend cannot pull `minio/minio` from Docker Hub any more; `SplitClusterFaultToleranceIT.testStreamJobCancelResolvesWhenWorkerCrashesBeforeCancelAck` and `RocketMqIT` also fail on the latest completed `dev` build; `kafka-connector-it` was cancelled by the job time limit. CI for `ea03511df99` is queued in the fork. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
