peterxcli opened a new pull request, #11343:
URL: https://github.com/apache/ozone/pull/11343

   ## What changes were proposed in this pull request?
   
   SCM removes a datanode from its NetworkTopology when the datanode becomes 
DEAD, and adds it back when the
   datanode starts heartbeating again and moves to HEALTHY_READONLY. Until now 
two event handlers did this:
   DeadNodeHandler removed the node and HealthyReadOnlyNodeHandler added it 
back. They run on separate
   EventQueue threads and act on a node state they read a little earlier, so 
their updates can land in the
   wrong order. HDDS-14834 narrowed the window with an extra state check in 
DeadNodeHandler, but this can
   still happen:
   
   1. DeadNodeHandler checks the node and sees it is still DEAD.
   2. The datanode heartbeats again. SCM moves it to HEALTHY_READONLY and fires 
HEALTHY_READONLY_NODE.
   3. HealthyReadOnlyNodeHandler adds the node to the topology. Nothing 
changes, because the node is still there.
   4. DeadNodeHandler removes the node.
   
   The datanode is then healthy but missing from the topology, so placement 
policies that pick nodes from the
   topology won't choose it.
   
   This PR moves the topology update into NodeStateManager, next to the health 
state change itself. The node is
   removed when it becomes DEAD and added back when it recovers. Both 
transitions happen inside the synchronized
   health check, so the topology always changes in the same order as the node's 
health. The handlers no longer
   touch the topology. That also removes the opposite race, where a late add 
puts back a node that has died
   again, and the parent-pointer checks in both handlers that could fail during 
the race. A failure while
   updating the topology is logged rather than thrown, because an exception 
there would stop SCM from
   scheduling further health checks.
   
   Test changes:
   - TestNodeStateManager: new tests for the topology updates when a node 
becomes DEAD and recovers, and for a
     failing topology update not breaking the health check.
   - TestDeadNodeHandler: removes the HDDS-14834 tests for handler behaviour 
that no longer exists. The
     remaining one now checks that DeadNodeHandler skips a node that is no 
longer DEAD.
   - TestQueryNode: new end-to-end test that stops a datanode, waits for it to 
leave SCM's topology, then
     restarts it and checks it is back.
   
   ## What is the link to the Apache JIRA
   
   https://issues.apache.org/jira/browse/HDDS-16623
   
   ## How was this patch tested?
   
   - All tests in the SCM `org.apache.hadoop.hdds.scm.node` package pass, except
     `TestSCMNodeManager#testScmClusterIsInExpectedState2`, which is 
timing-sensitive and also fails
     intermittently on master (2 of 3 local runs on unmodified master).
   - `TestQueryNode` in `ozone-integration-test`, including the new end-to-end 
test.
   - checkstyle and PMD.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to