[
https://issues.apache.org/jira/browse/HADOOP-19986?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18112432#comment-18112432
]
ASF GitHub Bot commented on HADOOP-19986:
-----------------------------------------
joseluisll opened a new pull request, #8723:
URL: https://github.com/apache/hadoop/pull/8723
### Description of PR
https://issues.apache.org/jira/browse/HADOOP-19986
`MiniDFSCluster.shutdownDataNodes()` tears DataNodes down strictly serially
- stop one, join it, move to the next. The DataNodes not yet reached keep
retrying a NameNode the test has already killed.
DataNodes in a single JVM share an `ipc.Client` through `ClientCache`, and
therefore share its per-address `Connection` objects. A surviving DataNode's
`BPServiceActor` holds a `Connection` monitor across its connect-retry sleeps
in `handleConnectionFailure`, while an actor of the DataNode being joined sits
BLOCKED on that same monitor in `Client.addCall`. `Thread.interrupt()` cannot
dislodge a BLOCKED thread, so the `stop()` issued by
`BlockPoolManager.shutDownAll` is ineffective and the join waits on scheduling
luck: the holder releases and re-acquires roughly every 2s, and unfair monitors
starve the blocked thread.
Measured on a 2-core runner: DataNode 2 took 165s to shut down, after which
DN1 and DN0 finished in ~10ms. That overran the 180s timeout of
`TestBalancerWithHANameNodes#testBalancerWithObserverWithFailedNode`, and runs
that stayed under the deadline still burned ~115s in teardown.
The fix signals every DataNode before joining any, so the monitor holder
aborts its sleep and releases it:
1. `BlockPoolManager#signalShutDownAll` - the stop-without-join half of the
existing `shutDownAll`, which now delegates to it.
2. `DataNode#signalBlockPoolShutdown` - `@VisibleForTesting`, null-safe,
signals every block pool service without waiting.
3. `MiniDFSCluster#shutdownDataNodes` - signals all DataNodes up front, then
runs the existing per-DataNode shutdown loop unchanged.
`stop()` is idempotent, so the per-DataNode shutdown path behaves exactly as
before; the only change is that the interrupts now all land before the first
join.
### How was this patch tested?
20 runs of
`TestBalancerWithHANameNodes#testBalancerWithObserverWithFailedNode` on a
2-core runner. Before: 4 of 20 anomalous (182.3s, 181.3s, 116.7s, 113.9s).
After: 20 of 20 passed within 44-48s.
Full CI on the fork, all jobs green - `common`, `hdfs - other`, `hdfs -
slow`, `hdfs-rbf`, `mr`, `other`, `yarn-server-rm` on Java 17, plus build-only
on Java 21 and Java 25:
https://github.com/joseluisll/hadoop/actions/runs/34147776139
That run was on commit `1700e04b`; this branch has since been rebased onto
current trunk with no change to the patch itself.
### For code changes:
- [x] Does the title of this PR start with the corresponding JIRA issue id
(e.g. 'HADOOP-17799. Your PR title ...')?
- [ ] Object storage: Have the integration tests been executed and the
endpoint declared according to the connector-specific documentation?
- [ ] If adding new dependencies to the code, are these dependencies
licensed in a way that is compatible for inclusion under [ASF
2.0](http://www.apache.org/legal/resolved.html#category-a)?
- [ ] If applicable, have you updated the `LICENSE`, `LICENSE-binary`,
`NOTICE-binary` files?
### AI Tooling
Contains content generated by Claude Code.
> MiniDFSCluster.shutdownDataNodes() should signal all DataNodes before joining
> any
> ---------------------------------------------------------------------------------
>
> Key: HADOOP-19986
> URL: https://issues.apache.org/jira/browse/HADOOP-19986
> Project: Hadoop Common
> Issue Type: Sub-task
> Components: hdfs, test
> Reporter: Jose Luis López
> Priority: Minor
>
> {{TestBalancerWithHANameNodes#testBalancerWithObserverWithFailedNode}}
> intermittently
> times out after 180s. The hang is in teardown, not in the balancer.
> h3. Cause
> {{MiniDFSCluster.shutdownDataNodes()}} tears DataNodes down strictly serially
> –
> stop one, join it, move to the next. The DataNodes not yet reached keep
> retrying
> the NameNode the test has already killed.
> DataNodes in a single JVM share an {{ipc.Client}} through
> {{{}ClientCache{}}}, and so
> share its per-address {{Connection}} objects. A surviving DataNode's
> {{BPServiceActor}} holds a {{Connection}} monitor across its connect-retry
> sleeps
> in {{{}handleConnectionFailure{}}}, while an actor of the DataNode being
> joined sits
> BLOCKED on that same monitor in {{{}Client.addCall{}}}.
> {{Thread.interrupt()}} cannot
> dislodge a BLOCKED thread, so the {{stop()}} issued by
> {{BlockPoolManager.shutDownAll}} is ineffective and the join waits on
> scheduling
> luck: the holder releases and re-acquires roughly every 2s, and unfair
> monitors
> starve the blocked thread.
> Measured on a 2-core runner: DataNode 2 took 165s to shut down, after which
> DN1 and
> DN0 finished in ~10ms.
> h3. Fix
> Signal every DataNode before joining any of them, so the monitor holder
> aborts its
> sleep and releases. Adds {{BlockPoolManager#signalShutDownAll}} (the
> stop-without-join
> half of the existing {{{}shutDownAll{}}}) and
> {{{}DataNode#signalBlockPoolShutdown{}}}.
> {{stop()}} is idempotent, so the existing per-DataNode shutdown path is
> unchanged.
> h3. Verification
> several (at least 3) 20 runs of the test on a 2-core GitHub runner, before
> and after:
> || ||before||after||
> |failures|4/20|0/20|
> |wall-clock spread|53s - 182s|44s - 48s|
> |test method time|up to 181s|14.4s - 15.0s|
> The two sub-timeout outliers before the fix (116.7s, 113.9s) are the same
> starvation landing under the deadline, so the test was silently burning ~2
> minutes
> on runs that reported as passing.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]