Chu Cheng Li created HDDS-16623:
-----------------------------------
Summary: Datanode can be missing from NetworkTopology after it
recovers from DEAD
Key: HDDS-16623
URL: https://issues.apache.org/jira/browse/HDDS-16623
Project: Apache Ozone
Issue Type: Improvement
Reporter: Chu Cheng Li
SCM removes a datanode from its NetworkTopology when the datanode becomes DEAD,
and adds it back when the datanode starts heartbeating again and moves to
HEALTHY_READONLY. Since HDDS-5437 this has been done by two event handlers:
DeadNodeHandler removes the node and HealthyReadOnlyNodeHandler adds it back.
The handlers run on separate EventQueue threads, and each one acts on a node
state it read a little earlier.
HDDS-14834 made DeadNodeHandler check the node state again just before removing
the node. The check and the removal are still separate steps, though, so this
can happen:
1. DeadNodeHandler checks the node and sees it is still DEAD.
2. The datanode heartbeats again. SCM moves it to HEALTHY_READONLY and fires
HEALTHY_READONLY_NODE.
3. HealthyReadOnlyNodeHandler adds the node to the topology. Nothing changes,
because the node is still there.
4. DeadNodeHandler removes the node.
The datanode is now HEALTHY_READONLY, and soon HEALTHY, but it is missing from
the topology. Placement policies that pick nodes from the topology won't choose
it until it next moves to HEALTHY_READONLY.
*Proposed fix*: update the topology in NodeStateManager at the same time as the
health state changes. Remove the node when it becomes DEAD, and add it back
when it recovers. Both changes happen inside the synchronized health check, so
the topology always changes in the same order as the node's health. The event
handlers no
longer need to touch it.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]