Sadanand Shenoy created HDDS-15962:
--------------------------------------

             Summary: Remove leader readiness check on the bootstrap flow
                 Key: HDDS-15962
                 URL: https://issues.apache.org/jira/browse/HDDS-15962
             Project: Apache Ozone
          Issue Type: Bug
            Reporter: Sadanand Shenoy
            Assignee: Sadanand Shenoy


Current logic rejects checkpoint requests with 503 unless {{isLeaderReady()}} 
is true.
In a degraded cluster (e.g., one OM down, one follower far behind), this 
creates a loop:
 # Lagging follower needs checkpoint to catch up.
 # Follower requests {{/v2/dbCheckpoint}} from leader.
 # Leader is {{LEADER_AND_NOT_READY}} and returns 503 due to 
{{isLeaderReady()}} check.
 # Follower cannot catch up, so leader readiness does not progress.
 # System remains stuck until another OM is brought back.

Checkpoint creation already has consistency barriers:

- awaitDoubleBufferFlush()
- bootstrap write lock (BOOTSTRAP_LOCK)
So requiring LEADER_AND_READY here is stricter than necessary and can block 
recovery. instead only reject if NOT_LEADER



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to