DanielLeens commented on PR #12284:
URL: https://github.com/apache/seatunnel/pull/12284#issuecomment-5662049411

   Follow-up to the second review round and to the fork's Build run for 
`e7bccdfb0b2` (run 34740882765). Commit `ea03511df99` addresses everything that 
is this PR's to fix.
   
   **Blocker 1, staged hand-off behind a column-appending wrapper: fixed.** The 
hand-off no longer requires positional equality between the hint-derived schema 
and the upstream produced table. `SQLLineageSchema.inOrderOf` adopts the 
upstream layout when it holds exactly the columns the hints produce (same names 
and definitions), keeping identities, creating hints and repositioning flags; 
anything else still fails fast with `TRANSFORM_COMMON-09`. For `Metadata -> Sql 
(select *)` the translator now emits `ADD age AFTER weight` (the last base 
column), which replays onto the pre-event output, so 
`ProductionPipelineSchemaChangeTest` passes by construction. 
`TransformChainLiveAlterTest` drives the transforms without the engine hand-off 
and keeps its tail positions. New unit tests: 
`testStagedHandOffAdoptsLayoutOfColumnAppendingWrapper` (Metadata upstream, 
rows through) and `testStagedHandOffRejectsLayoutWithOtherColumns`. The error 
message names the event class and type when the event carr
 ies no statement instead of printing `'null'`. Docs describe the adopted 
layout.
   
   **Blocker 2, `TestFilterRowKindIT.testFilterRowKindMultiTable`: not caused 
by this PR, but checked rather than dismissed.** Unmatched tables of that job 
do go through `IdentityMapTransform`, which this PR touches. In the same run 
the job passed on Zeta, Spark 2.4.6, Spark 3.3.0, Flink 1.13.6 and Flink 1.18.0 
and failed only on the Flink 1.15.3 container, where the log shows the Assert 
sink failing `MIN_ROW 100` with 25 and then 91 rows, the Flink job restarting 
twice, and then `MAX_ROW 100` failing with 150 rows: counts accumulated across 
restarts, the `AssertSinkWriter` static-counter problem that has its own fix in 
flight. An identity regression would fail on every engine.
   
   **This PR's own E2E failures: root cause in `MultiTableSinkWriter`, worked 
around and documented.** With `multi_table_sink_replica = 2` the writer 
dispatches every schema change to both sub-writers of the same physical table. 
The second sub-writer replays `CHANGE reuse_col reuse_col_old, ADD reuse_col` 
after the first one already re-added `reuse_col`, so MySQL answers `Duplicate 
column name 'reuse_col_old'` and the job restarts forever; `DROP name, ADD 
name` is executed twice, and the projection test's query happened to run 
between the two statements, which aborted the test instead of retrying. The 
same events reach the sink without any transform, so this is independent of the 
translator; the physical DDL should be applied once per physical table and 
siblings refreshed at runtime only, the way `RestoreTableSchemaEvent` already 
is. Recorded as a follow-up in the description. The two SQL jobs now run with 
one sub-writer per table (the pre-existing jobs keep replica 2), and polls tha
 t hit the moment between the two statements of one composite are retried.
   
   **Remaining failures in that run, none in this PR's scope:** Iceberg S3, S3 
file, Paimon S3 and Databend cannot pull `minio/minio` from Docker Hub any 
more; 
`SplitClusterFaultToleranceIT.testStreamJobCancelResolvesWhenWorkerCrashesBeforeCancelAck`
 and `RocketMqIT` also fail on the latest completed `dev` build; 
`kafka-connector-it` was cancelled by the job time limit.
   
   CI for `ea03511df99` is queued in the fork.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to