[ 
https://issues.apache.org/jira/browse/HADOOP-19986?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18112432#comment-18112432
 ] 

ASF GitHub Bot commented on HADOOP-19986:
-----------------------------------------

joseluisll opened a new pull request, #8723:
URL: https://github.com/apache/hadoop/pull/8723

   ### Description of PR
   
   https://issues.apache.org/jira/browse/HADOOP-19986
   
   `MiniDFSCluster.shutdownDataNodes()` tears DataNodes down strictly serially 
- stop one, join it, move to the next. The DataNodes not yet reached keep 
retrying a NameNode the test has already killed.
   
   DataNodes in a single JVM share an `ipc.Client` through `ClientCache`, and 
therefore share its per-address `Connection` objects. A surviving DataNode's 
`BPServiceActor` holds a `Connection` monitor across its connect-retry sleeps 
in `handleConnectionFailure`, while an actor of the DataNode being joined sits 
BLOCKED on that same monitor in `Client.addCall`. `Thread.interrupt()` cannot 
dislodge a BLOCKED thread, so the `stop()` issued by 
`BlockPoolManager.shutDownAll` is ineffective and the join waits on scheduling 
luck: the holder releases and re-acquires roughly every 2s, and unfair monitors 
starve the blocked thread.
   
   Measured on a 2-core runner: DataNode 2 took 165s to shut down, after which 
DN1 and DN0 finished in ~10ms. That overran the 180s timeout of 
`TestBalancerWithHANameNodes#testBalancerWithObserverWithFailedNode`, and runs 
that stayed under the deadline still burned ~115s in teardown.
   
   The fix signals every DataNode before joining any, so the monitor holder 
aborts its sleep and releases it:
   
   1. `BlockPoolManager#signalShutDownAll` - the stop-without-join half of the 
existing `shutDownAll`, which now delegates to it.
   2. `DataNode#signalBlockPoolShutdown` - `@VisibleForTesting`, null-safe, 
signals every block pool service without waiting.
   3. `MiniDFSCluster#shutdownDataNodes` - signals all DataNodes up front, then 
runs the existing per-DataNode shutdown loop unchanged.
   
   `stop()` is idempotent, so the per-DataNode shutdown path behaves exactly as 
before; the only change is that the interrupts now all land before the first 
join.
   
   ### How was this patch tested?
   
   20 runs of 
`TestBalancerWithHANameNodes#testBalancerWithObserverWithFailedNode` on a 
2-core runner. Before: 4 of 20 anomalous (182.3s, 181.3s, 116.7s, 113.9s). 
After: 20 of 20 passed within 44-48s.
   
   Full CI on the fork, all jobs green - `common`, `hdfs - other`, `hdfs - 
slow`, `hdfs-rbf`, `mr`, `other`, `yarn-server-rm` on Java 17, plus build-only 
on Java 21 and Java 25:
   https://github.com/joseluisll/hadoop/actions/runs/34147776139
   
   That run was on commit `1700e04b`; this branch has since been rebased onto 
current trunk with no change to the patch itself.
   
   ### For code changes:
   
   - [x] Does the title of this PR start with the corresponding JIRA issue id 
(e.g. 'HADOOP-17799. Your PR title ...')?
   - [ ] Object storage: Have the integration tests been executed and the 
endpoint declared according to the connector-specific documentation?
   - [ ] If adding new dependencies to the code, are these dependencies 
licensed in a way that is compatible for inclusion under [ASF 
2.0](http://www.apache.org/legal/resolved.html#category-a)?
   - [ ] If applicable, have you updated the `LICENSE`, `LICENSE-binary`, 
`NOTICE-binary` files?
   
   ### AI Tooling
   
   Contains content generated by Claude Code.
   




> MiniDFSCluster.shutdownDataNodes() should signal all DataNodes before joining 
> any
> ---------------------------------------------------------------------------------
>
>                 Key: HADOOP-19986
>                 URL: https://issues.apache.org/jira/browse/HADOOP-19986
>             Project: Hadoop Common
>          Issue Type: Sub-task
>          Components: hdfs, test
>            Reporter: Jose Luis López
>            Priority: Minor
>
> {{TestBalancerWithHANameNodes#testBalancerWithObserverWithFailedNode}} 
> intermittently
> times out after 180s. The hang is in teardown, not in the balancer.
> h3. Cause
> {{MiniDFSCluster.shutdownDataNodes()}} tears DataNodes down strictly serially 
> –
> stop one, join it, move to the next. The DataNodes not yet reached keep 
> retrying
> the NameNode the test has already killed.
> DataNodes in a single JVM share an {{ipc.Client}} through 
> {{{}ClientCache{}}}, and so
> share its per-address {{Connection}} objects. A surviving DataNode's
> {{BPServiceActor}} holds a {{Connection}} monitor across its connect-retry 
> sleeps
> in {{{}handleConnectionFailure{}}}, while an actor of the DataNode being 
> joined sits
> BLOCKED on that same monitor in {{{}Client.addCall{}}}. 
> {{Thread.interrupt()}} cannot
> dislodge a BLOCKED thread, so the {{stop()}} issued by
> {{BlockPoolManager.shutDownAll}} is ineffective and the join waits on 
> scheduling
> luck: the holder releases and re-acquires roughly every 2s, and unfair 
> monitors
> starve the blocked thread.
> Measured on a 2-core runner: DataNode 2 took 165s to shut down, after which 
> DN1 and
> DN0 finished in ~10ms.
> h3. Fix
> Signal every DataNode before joining any of them, so the monitor holder 
> aborts its
> sleep and releases. Adds {{BlockPoolManager#signalShutDownAll}} (the 
> stop-without-join
> half of the existing {{{}shutDownAll{}}}) and 
> {{{}DataNode#signalBlockPoolShutdown{}}}.
> {{stop()}} is idempotent, so the existing per-DataNode shutdown path is 
> unchanged.
> h3. Verification
> several (at least 3) 20 runs of the test on a 2-core GitHub runner, before 
> and after:
> || ||before||after||
> |failures|4/20|0/20|
> |wall-clock spread|53s - 182s|44s - 48s|
> |test method time|up to 181s|14.4s - 15.0s|
> The two sub-timeout outliers before the fix (116.7s, 113.9s) are the same
> starvation landing under the deadline, so the test was silently burning ~2 
> minutes
> on runs that reported as passing.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to