[ 
https://issues.apache.org/jira/browse/KAFKA-20785?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Justine Olshan resolved KAFKA-20785.
------------------------------------
    Resolution: Fixed

> Unexpected InvalidTxnStateException on EndTxn when commit races with 
> transaction timeout
> ----------------------------------------------------------------------------------------
>
>                 Key: KAFKA-20785
>                 URL: https://issues.apache.org/jira/browse/KAFKA-20785
>             Project: Kafka
>          Issue Type: Bug
>          Components: clients, producer , streams
>            Reporter: Lianet Magrans
>            Assignee: Gavin Wang
>            Priority: Major
>
> We are seeing cases where a Kafka Streams thread stalls briefly, and during 
> the stall the broker correctly aborts the open transaction (triggered by the 
> transaction.timeout.ms).
> However, the producer was already in COMMITTING_TRANSACTION when the stall 
> began, so when the client resumes, its stale EndTxn(commit) reaches the 
> coordinator. Under transactions V2 the retry-tolerant epoch check accepts the 
> request (the timeout abort bumped the epoch, so stored epoch == request epoch 
> + 1 matches the retry condition), and it reaches the state check instead: the 
> coordinator logs "TransactionalId ... state is COMPLETE_ABORT, but received 
> transaction marker result to send: COMMIT" and returns INVALID_TXN_STATE.
> The client treats INVALID_TXN_STATE on EndTxn as unconditionally fatal, so 
> the producer transitions to FATAL_ERROR with InvalidTxnStateException. Kafka 
> Streams does not recognize this as recoverable and the StreamThread dies.
> The transaction was already fully rolled back server side, so the client did 
> nothing wrong — the same race under transaction V1 returns PRODUCER_FENCED 
> (the strict epoch check fails first), which Streams handles gracefully as a 
> rebalance. The fatal outcome is V2 only.
> This gap on the EndTxn path seems the same as what was reported/fix on the 
> Produce path with KAFKA-19690: producers that are mid-produce when the tx 
> times out. The fix was to return a recoverable InvalidProducerEpochException 
> instead in this case. 
> In this case of EndTxn getting to the broker when the tx was already aborted 
> on timeout, since the commit is guaranteed not to have taken effect, the 
> client should be able to abort and continue rather than fail fatally.
> Seems we have options to address here: fix on the client (treat 
> InvalidTxState on EndTxn as abortable so the app can abort + retry), and/or 
> on the broker (change error returned for this case? analogous to 
> KAFKA-19690's approach on the produce path).



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to