joseluisll opened a new pull request, #8723: URL: https://github.com/apache/hadoop/pull/8723
### Description of PR https://issues.apache.org/jira/browse/HADOOP-19986 `MiniDFSCluster.shutdownDataNodes()` tears DataNodes down strictly serially - stop one, join it, move to the next. The DataNodes not yet reached keep retrying a NameNode the test has already killed. DataNodes in a single JVM share an `ipc.Client` through `ClientCache`, and therefore share its per-address `Connection` objects. A surviving DataNode's `BPServiceActor` holds a `Connection` monitor across its connect-retry sleeps in `handleConnectionFailure`, while an actor of the DataNode being joined sits BLOCKED on that same monitor in `Client.addCall`. `Thread.interrupt()` cannot dislodge a BLOCKED thread, so the `stop()` issued by `BlockPoolManager.shutDownAll` is ineffective and the join waits on scheduling luck: the holder releases and re-acquires roughly every 2s, and unfair monitors starve the blocked thread. Measured on a 2-core runner: DataNode 2 took 165s to shut down, after which DN1 and DN0 finished in ~10ms. That overran the 180s timeout of `TestBalancerWithHANameNodes#testBalancerWithObserverWithFailedNode`, and runs that stayed under the deadline still burned ~115s in teardown. The fix signals every DataNode before joining any, so the monitor holder aborts its sleep and releases it: 1. `BlockPoolManager#signalShutDownAll` - the stop-without-join half of the existing `shutDownAll`, which now delegates to it. 2. `DataNode#signalBlockPoolShutdown` - `@VisibleForTesting`, null-safe, signals every block pool service without waiting. 3. `MiniDFSCluster#shutdownDataNodes` - signals all DataNodes up front, then runs the existing per-DataNode shutdown loop unchanged. `stop()` is idempotent, so the per-DataNode shutdown path behaves exactly as before; the only change is that the interrupts now all land before the first join. ### How was this patch tested? 20 runs of `TestBalancerWithHANameNodes#testBalancerWithObserverWithFailedNode` on a 2-core runner. Before: 4 of 20 anomalous (182.3s, 181.3s, 116.7s, 113.9s). After: 20 of 20 passed within 44-48s. Full CI on the fork, all jobs green - `common`, `hdfs - other`, `hdfs - slow`, `hdfs-rbf`, `mr`, `other`, `yarn-server-rm` on Java 17, plus build-only on Java 21 and Java 25: https://github.com/joseluisll/hadoop/actions/runs/34147776139 That run was on commit `1700e04b`; this branch has since been rebased onto current trunk with no change to the patch itself. ### For code changes: - [x] Does the title of this PR start with the corresponding JIRA issue id (e.g. 'HADOOP-17799. Your PR title ...')? - [ ] Object storage: Have the integration tests been executed and the endpoint declared according to the connector-specific documentation? - [ ] If adding new dependencies to the code, are these dependencies licensed in a way that is compatible for inclusion under [ASF 2.0](http://www.apache.org/legal/resolved.html#category-a)? - [ ] If applicable, have you updated the `LICENSE`, `LICENSE-binary`, `NOTICE-binary` files? ### AI Tooling Contains content generated by Claude Code. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
