Rangsh commented on PR #12081: URL: https://github.com/apache/seatunnel/pull/12081#issuecomment-5979215812
@SEZ9 — re-dispatched as requested. Failed jobs re-run on the same head `1cbac0086` with no code changes: [Build 36520988425, attempt 2](https://github.com/Rangsh/seatunnel/actions/runs/36520988425). **The three ITs you asked about** | IT | JDK 8 | JDK 11 | |---|---|---| | `BackpressureSlowSinkIT` | reproduced (`only observed 2` of the expected 3 checkpoints) | reproduced (same assertion) | | `CheckpointCoordinatorFailoverIT#testBatchJobCompletesAfterMasterFailoverDuringCloseHandshake` | reproduced (`expected: <610> but was: <620>`) | reproduced (same) | | `SavepointBusySourceBarrierIT` | passed | passed | As you suggested, I checked the surefire output for the two that reproduced. Within each test's log section there are zero occurrences of `TimeoutException`, `RequestFuture`, `queryExecuteStatus`, `batchQueryExecuteFailsStatus`, `WALWorkHandler` or `IMapStorageException`, so by your criterion the WAL/checkpoint-storage path does not show up in either of them. Both ITs came into this branch through the `dev` sync ([#12107](https://github.com/apache/seatunnel/pull/12107), [#12031](https://github.com/apache/seatunnel/pull/12031)). Note that they now failed on both JDKs in both attempts. `all-connectors-it-1`, `all-connectors-it-7` and `doris-connector-it` all passed on the re-run. **One new failure that does touch the WAL path, flagged for transparency** `SplitClusterFaultToleranceIT#testStreamJobRestoreInAllNodeDown` failed on JDK 11 only (passed on JDK 8): `expected: <CANCELED> but was: <UNKNOWABLE>`. Timeline from the log: 1. `08:27:02` the test cancels the job. 2. One IMap write-through store (`requestId 701`) waits the full `60000 ms`. The new F8 WARN fires at `08:28:02.580`: `wait for write status timed out for requestId 701 after 60000 ms (limit 60000 ms)`. 3. `FileMapStore.store()` then throws `IMapStorageException: IMap store failed to persist durably`, the source task ends FAILED, and the job resolves to `UNKNOWABLE` instead of `CANCELED`. 4. 16 ms later the WAL worker completes the write: `requestId is 701 not found in RequestFutureCache`. The 60 s stall itself is not new: the pre-PR timed `get` also waited the full timeout. But before this PR `FileMapStore.store()` discarded the `false` and the job carried on. With the fail-loud propagation added in this PR (part of the F4 resolution), the same stall now fails the task. So I can't call this one unrelated. My reading is that the PR surfaces a pre-existing slow-write condition rather than causing it, but it does change the observable outcome of this test. How would you like to handle it? Options I see: (a) keep the fail-loud behaviour as intended and treat the slow write during cancellation as a separate issue, or (b) adjust how a store timeout is handled during cancellation in this PR. Happy to go either way. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
