[
https://issues.apache.org/jira/browse/HADOOP-19986?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18112497#comment-18112497
]
ASF GitHub Bot commented on HADOOP-19986:
-----------------------------------------
joseluisll commented on PR #8723:
URL: https://github.com/apache/hadoop/pull/8723#issuecomment-5579024588
@pan3793 @slfan1989 This PR is Green and ready to be reviewed. This one
solves race conditions on minicluster when asked to be shutdown, it hanged on
the closing datanodes. This surfaced because up to a month ago minicluster was
not shutdown in the tests, leaked. I worked in the jiras to fix that. Properly
closing the miniclusters and eliminating this race condition improves CI
testing.
It is a prerrequisite for HADOOP-19979, that will fix 4 flaky conditions so
that we reduce later the GHA exclude list, as explained on HADOOP-19981
[umbrella for all GHA reactivation candidates from excluded-tests.txt] and its
subtasks [specific identified candidates for reactivation].
> MiniDFSCluster.shutdownDataNodes() should signal all DataNodes before joining
> any
> ---------------------------------------------------------------------------------
>
> Key: HADOOP-19986
> URL: https://issues.apache.org/jira/browse/HADOOP-19986
> Project: Hadoop Common
> Issue Type: Sub-task
> Components: hdfs, test
> Reporter: Jose Luis López
> Priority: Minor
> Labels: pull-request-available
>
> {{TestBalancerWithHANameNodes#testBalancerWithObserverWithFailedNode}}
> intermittently
> times out after 180s. The hang is in teardown, not in the balancer.
> h3. Cause
> {{MiniDFSCluster.shutdownDataNodes()}} tears DataNodes down strictly serially
> –
> stop one, join it, move to the next. The DataNodes not yet reached keep
> retrying
> the NameNode the test has already killed.
> DataNodes in a single JVM share an {{ipc.Client}} through
> {{{}ClientCache{}}}, and so
> share its per-address {{Connection}} objects. A surviving DataNode's
> {{BPServiceActor}} holds a {{Connection}} monitor across its connect-retry
> sleeps
> in {{{}handleConnectionFailure{}}}, while an actor of the DataNode being
> joined sits
> BLOCKED on that same monitor in {{{}Client.addCall{}}}.
> {{Thread.interrupt()}} cannot
> dislodge a BLOCKED thread, so the {{stop()}} issued by
> {{BlockPoolManager.shutDownAll}} is ineffective and the join waits on
> scheduling
> luck: the holder releases and re-acquires roughly every 2s, and unfair
> monitors
> starve the blocked thread.
> Measured on a 2-core runner: DataNode 2 took 165s to shut down, after which
> DN1 and
> DN0 finished in ~10ms.
> h3. Fix
> Signal every DataNode before joining any of them, so the monitor holder
> aborts its
> sleep and releases. Adds {{BlockPoolManager#signalShutDownAll}} (the
> stop-without-join
> half of the existing {{{}shutDownAll{}}}) and
> {{{}DataNode#signalBlockPoolShutdown{}}}.
> {{stop()}} is idempotent, so the existing per-DataNode shutdown path is
> unchanged.
> h3. Verification
> several (at least 3) 20 runs of the test on a 2-core GitHub runner, before
> and after:
> || ||before||after||
> |failures|4/20|0/20|
> |wall-clock spread|53s - 182s|44s - 48s|
> |test method time|up to 181s|14.4s - 15.0s|
> The two sub-timeout outliers before the fix (116.7s, 113.9s) are the same
> starvation landing under the deadline, so the test was silently burning ~2
> minutes
> on runs that reported as passing.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]