Rangsh commented on PR #12081:
URL: https://github.com/apache/seatunnel/pull/12081#issuecomment-5979215812

   @SEZ9 — re-dispatched as requested. Failed jobs re-run on the same head 
`1cbac0086` with no code changes: [Build 36520988425, attempt 
2](https://github.com/Rangsh/seatunnel/actions/runs/36520988425).
   
   **The three ITs you asked about**
   
   | IT | JDK 8 | JDK 11 |
   |---|---|---|
   | `BackpressureSlowSinkIT` | reproduced (`only observed 2` of the expected 3 
checkpoints) | reproduced (same assertion) |
   | 
`CheckpointCoordinatorFailoverIT#testBatchJobCompletesAfterMasterFailoverDuringCloseHandshake`
 | reproduced (`expected: <610> but was: <620>`) | reproduced (same) |
   | `SavepointBusySourceBarrierIT` | passed | passed |
   
   As you suggested, I checked the surefire output for the two that reproduced. 
Within each test's log section there are zero occurrences of 
`TimeoutException`, `RequestFuture`, `queryExecuteStatus`, 
`batchQueryExecuteFailsStatus`, `WALWorkHandler` or `IMapStorageException`, so 
by your criterion the WAL/checkpoint-storage path does not show up in either of 
them. Both ITs came into this branch through the `dev` sync 
([#12107](https://github.com/apache/seatunnel/pull/12107), 
[#12031](https://github.com/apache/seatunnel/pull/12031)). Note that they now 
failed on both JDKs in both attempts.
   
   `all-connectors-it-1`, `all-connectors-it-7` and `doris-connector-it` all 
passed on the re-run.
   
   **One new failure that does touch the WAL path, flagged for transparency**
   
   `SplitClusterFaultToleranceIT#testStreamJobRestoreInAllNodeDown` failed on 
JDK 11 only (passed on JDK 8): `expected: <CANCELED> but was: <UNKNOWABLE>`. 
Timeline from the log:
   
   1. `08:27:02` the test cancels the job.
   2. One IMap write-through store (`requestId 701`) waits the full `60000 ms`. 
The new F8 WARN fires at `08:28:02.580`: `wait for write status timed out for 
requestId 701 after 60000 ms (limit 60000 ms)`.
   3. `FileMapStore.store()` then throws `IMapStorageException: IMap store 
failed to persist durably`, the source task ends FAILED, and the job resolves 
to `UNKNOWABLE` instead of `CANCELED`.
   4. 16 ms later the WAL worker completes the write: `requestId is 701 not 
found in RequestFutureCache`.
   
   The 60 s stall itself is not new: the pre-PR timed `get` also waited the 
full timeout. But before this PR `FileMapStore.store()` discarded the `false` 
and the job carried on. With the fail-loud propagation added in this PR (part 
of the F4 resolution), the same stall now fails the task. So I can't call this 
one unrelated. My reading is that the PR surfaces a pre-existing slow-write 
condition rather than causing it, but it does change the observable outcome of 
this test.
   
   How would you like to handle it? Options I see: (a) keep the fail-loud 
behaviour as intended and treat the slow write during cancellation as a 
separate issue, or (b) adjust how a store timeout is handled during 
cancellation in this PR. Happy to go either way.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to