surafel58 commented on PR #11827:
URL: https://github.com/apache/seatunnel/pull/11827#issuecomment-5366357668

   Thanks @SEZ9. I have taken the documentation half of both findings on this 
PR, and consolidated the two behavioral changes into the follow-up #11911.
   
   Docs (commit 2ce10e84, EN and ZH): added a "Checkpoint Flush and Failure 
Handling" section covering both points:
   - Replay safety depends on the receiver accepting an exact duplicate (same 
labels, timestamp, value); a receiver that rejects a 
same-timestamp-different-value or out-of-order sample (Prometheus TSDB, and 
Cortex/Mimir/Thanos, return 400) can fail the replayed flush.
   - Transient failures fail the checkpoint; note to raise the engine's 
tolerableCheckpointFailureNumber on Spark/Flink, with the bounded retry tracked 
in #11911.
   
   Behavioral fixes -> #11911: both remaining items are flush() error-handling 
changes, so I grouped them in the follow-up rather than expanding this PR: (a) 
bounded retry with backoff for retryable failures (5xx/429/transport), and (b) 
treat duplicate/out-of-order 400 rejections as delivered so a replay after 
restore does not loop the checkpoint. I have widened #11911 to cover both and 
will implement it right after this PR merges (both change flush(), so doing it 
here would conflict).
   
   F3 is resolved, and F1/F2 now have the documentation you named as the 
minimum, with the behavioral work scoped and tracked in #11911. Would you be 
comfortable merging this PR with the behavioral changes handled in that 
follow-up? Happy to bring any of it into this PR instead if you would prefer.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to