[ 
https://issues.apache.org/jira/browse/HDDS-15962?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

ASF GitHub Bot updated HDDS-15962:
----------------------------------
    Labels: pull-request-available  (was: )

> Remove leader readiness check on the bootstrap flow
> ---------------------------------------------------
>
>                 Key: HDDS-15962
>                 URL: https://issues.apache.org/jira/browse/HDDS-15962
>             Project: Apache Ozone
>          Issue Type: Bug
>            Reporter: Sadanand Shenoy
>            Assignee: Sadanand Shenoy
>            Priority: Major
>              Labels: pull-request-available
>
> Current logic rejects checkpoint requests with 503 unless {{isLeaderReady()}} 
> is true.
> In a degraded cluster (e.g., one OM down, one follower far behind), this 
> creates a loop:
>  # Lagging follower needs checkpoint to catch up.
>  # Follower requests {{/v2/dbCheckpoint}} from leader.
>  # Leader is {{LEADER_AND_NOT_READY}} and returns 503 due to 
> {{isLeaderReady()}} check.
>  # Follower cannot catch up, so leader readiness does not progress.
>  # System remains stuck until another OM is brought back.
> Checkpoint creation already has consistency barriers:
> - awaitDoubleBufferFlush()
> - bootstrap write lock (BOOTSTRAP_LOCK)
> So requiring LEADER_AND_READY here is stricter than necessary and can block 
> recovery. instead only reject if NOT_LEADER



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to