[
https://issues.apache.org/jira/browse/HDDS-15962?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated HDDS-15962:
----------------------------------
Labels: pull-request-available (was: )
> Remove leader readiness check on the bootstrap flow
> ---------------------------------------------------
>
> Key: HDDS-15962
> URL: https://issues.apache.org/jira/browse/HDDS-15962
> Project: Apache Ozone
> Issue Type: Bug
> Reporter: Sadanand Shenoy
> Assignee: Sadanand Shenoy
> Priority: Major
> Labels: pull-request-available
>
> Current logic rejects checkpoint requests with 503 unless {{isLeaderReady()}}
> is true.
> In a degraded cluster (e.g., one OM down, one follower far behind), this
> creates a loop:
> # Lagging follower needs checkpoint to catch up.
> # Follower requests {{/v2/dbCheckpoint}} from leader.
> # Leader is {{LEADER_AND_NOT_READY}} and returns 503 due to
> {{isLeaderReady()}} check.
> # Follower cannot catch up, so leader readiness does not progress.
> # System remains stuck until another OM is brought back.
> Checkpoint creation already has consistency barriers:
> - awaitDoubleBufferFlush()
> - bootstrap write lock (BOOTSTRAP_LOCK)
> So requiring LEADER_AND_READY here is stricter than necessary and can block
> recovery. instead only reject if NOT_LEADER
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]