[
https://issues.apache.org/jira/browse/RATIS-2661?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Yongzao Dan reassigned RATIS-2661:
----------------------------------
Assignee: Xinyu Tan
> Recovered Raft group with no persisted configuration has no path out of
> STARTING
> --------------------------------------------------------------------------------
>
> Key: RATIS-2661
> URL: https://issues.apache.org/jira/browse/RATIS-2661
> Project: Ratis
> Issue Type: Bug
> Components: raft-group, server
> Affects Versions: 3.2.2
> Environment: Apache Ratis 3.2.2, revision
> 288c032064ce3d168b8a763e248a326459a4a9b7. gRPC transport, three voting peers,
> dynamic multi-group usage. Observed in an IoTDB-based 2.0.10.2 deployment.
> Not yet verified against newer Ratis releases.
> Reporter: Yongzao Dan
> Assignee: Xinyu Tan
> Priority: Major
> Labels: liveness
> Attachments: ratis-3.2.2-uninitialized-group-recovery-sanitized.log
>
>
> h2. Summary
> A dynamically managed Raft group can become permanently unrecoverable if all
> peers restart before the initial configuration entry is persisted.
> During automatic directory recovery, Ratis reconstructs the group using only
> its group ID. If no configuration exists in storage, the recovered group has
> an empty peer list. The division remains in STARTING with reason NOT_IN_CONF.
> The embedding application still knows the original membership and invokes
> setConfiguration with SET_UNCONDITIONALLY, but RaftServerImpl rejects the
> request because the lifecycle is not RUNNING. Since a group with an empty
> configuration cannot elect a leader, there is no apparent supported path to
> restore the membership.
> h2. Observed failure sequence
> # A new three-peer group was created with the correct initial membership.
> # Peer RPC failures prevented the group from forming a majority.
> # The group remained at term 0 and reached PRE_VOTE round 7808.
> # No Raft log or configuration entry was created; shutdown reported stopIndex
> = -1.
> # All servers were restarted while preserving their storage directories.
> # Each server recovered the group with peers:[] and entered STARTING /
> FOLLOWER / NOT_IN_CONF.
> # The application invoked SET_UNCONDITIONALLY with the original three peers.
> # Every request failed with ServerNotReadyException because the group was
> still STARTING.
> # After peer communication recovered, newly created groups elected leaders
> normally, but this recovered group remained unavailable.
> A sanitized log excerpt containing this sequence is attached.
> h2. Suspected root cause
> In Ratis 3.2.2, RaftServerProxy.initGroupDir() recovers a directory as:
> {code:java}
> addGroup(RaftGroup.valueOf(groupId), StartupOption.RECOVER);
> {code}
> This RaftGroup contains no peers. ServerState.initialize() only replaces that
> empty configuration if readRaftConfiguration() returns a persisted
> configuration.
> RaftServerImpl.start() does not call startAsPeer() when the local peer is
> absent from the configuration. It sets the role to FOLLOWER with NOT_IN_CONF,
> while the lifecycle remains STARTING. startAsPeer() is the path that
> transitions the lifecycle to RUNNING.
> RaftServerImpl.setConfigurationAsync() then rejects the recovery request
> before inspecting its mode:
> {code:java}
> assertLifeCycleState(LifeCycle.States.RUNNING);
> {code}
> The normal configuration-change path also checks for a ready leader. This
> creates an unbreakable liveness cycle:
> * Membership is required to elect a leader and reach RUNNING.
> * RUNNING and a leader are required to restore membership.
> Simply relaxing the lifecycle assertion may therefore be insufficient; a
> guarded bootstrap path would also need to handle the absence of a leader.
> h2. Proposed reproduction
> This reproduction is inferred from the production incident and has not yet
> been reduced to a standalone Ratis test.
> # Start three Ratis servers and add a new group containing all three peers.
> # Block peer RPC before the first configuration entry is written or committed.
> # Verify that the group remains at term 0 with last log index -1.
> # Stop all servers while preserving their group directories.
> # Restart them through the StartupOption.RECOVER directory-scanning path.
> # Call setConfiguration with SET_UNCONDITIONALLY and the original peer list.
> # Verify that every division remains in STARTING and rejects the request.
> h2. Expected behavior
> Recovery should not leave an uninitialized group in a state with no supported
> transition to RUNNING. Ratis should provide a non-destructive way to restore
> bootstrap membership when no configuration or Raft log has ever been
> persisted.
> h2. Impact
> A transient bootstrap communication failure followed by a restart becomes a
> permanent outage for that group. Repeated restarts and configuration retries
> do not help. Recovery currently requires out-of-band storage intervention or
> application-specific group recreation.
> h2. Possible fix directions
> * Persist enough bootstrap membership information when the group directory is
> created.
> * Allow callers to provide membership for recovered group IDs that have no
> stored configuration.
> * Provide a guarded re-bootstrap operation for STARTING + empty configuration
> + empty log. If SET_UNCONDITIONALLY is reused, it would also need a safe
> bootstrap path that does not depend on an existing leader.
> Is there an existing supported recovery procedure for this state? If so,
> documenting that procedure would also help dynamic multi-group applications.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)