[
https://issues.apache.org/jira/browse/HADOOP-19979?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18112182#comment-18112182
]
ASF GitHub Bot commented on HADOOP-19979:
-----------------------------------------
joseluisll commented on PR #8717:
URL: https://github.com/apache/hadoop/pull/8717#issuecomment-5566244310
> Pre-existing, but the leak fix now relies on it:
`TimelineReaderServer.serviceStop()` calls `readerWebServer.stop()` before
`super.serviceStop()` with no try/finally. `HttpServer2.stop()` rethrows its
`MultiException`, and the service is already STOPPED by then so a retry is a
no-op, leaving the storage monitor running. `try { readerWebServer.stop(); }
finally { super.serviceStop(); }` closes it; fine as a follow-up if this PR
stays test-only.
>
> TestStandbyCheckpoints.java L844: `testPutFsimagePartFailed` has the same
shape: `snnCheckpointTime2` is read right after `waitForCheckpoint(cluster, 0,
[23])` and asserted `> snnCheckpointTime1`, but `nns[1]` stamps
`lastCheckpointTime` only after `doCheckpoint()` returns, which here is after
the upload to the stopped `nns[2]` fails. The same wait fixes it.
TestStandbyCheckpoints.testPutFsimagePartFailed — good catch, it's the same
race. Fixed here rather than deferred: nns[2] is shut down in that test, so
nns[1] is the uploader and the anchor is the same one used in
testLastCheckpointTime — waitForCheckpoint(cluster, 1, ImmutableList.of(23))
followed by a wait on getStandbyLastCheckpointTime() moving past the baseline.
TimelineReaderServer.serviceStop() — fixed here too, since the leak fix
depends on it: try { readerWebServer.stop(); } finally { super.serviceStop();
}, so a web server that fails to stop can't leave the storage monitor's polling
executor behind.
> Fix four tests that assert on work owned by another thread
> ----------------------------------------------------------
>
> Key: HADOOP-19979
> URL: https://issues.apache.org/jira/browse/HADOOP-19979
> Project: Hadoop Common
> Issue Type: Test
> Components: common, test
> Reporter: Jose Luis López
> Priority: Critical
> Labels: pull-request-available
>
> Four tests assert on, or tear down around, work owned by another thread
> without
> waiting for it or stopping it. All four fail intermittently, and none of the
> failures say anything about the code under test.
> * {{TestSSLHttpServerMTLS.testUntrustedClientIsRejected}} (common) expects an
> SSLHandshakeException, but the server's close races the client's last
> handshake
> flight; when the close wins the client gets a SocketException instead. 7 of 25
> runs fail. Assert that the request is refused rather than which exception
> carries it.
> * {{TestLogAggregationService.testLocalFileDeletionAfterUpload}}
> (nodemanager)
> waits for each log file to go, then asserts on the parent directory with no
> wait; DeletionService removes files before the directories holding them. Hit 4
> of the 60 most recent PRs, including unrelated ones. Wait for the directory
> too.
> * {{TestStandbyCheckpoints.testLastCheckpointTime}} (hdfs) waits for the
> active
> to hold the new image, then reads a standby's checkpoint time, which that wait
> does not cover: any standby may be the uploader, and it stamps
> lastCheckpointTime only after doCheckpoint() returns, so the interval can
> read 0
> against an expected 3000. Wait for that value to move.
> * {{TestTimelineReaderHBaseDown}} (timelineservice-hbase-tests) starts a
> TimelineReaderServer in all five tests and never stops one, leaking the
> TimelineStorageMonitor it schedules: non-daemon threads polling a minicluster
> the test has torn down. The module builds with forkCount 0, so these
> accumulate
> across its eleven test classes. Stop the server in a finally, as every other
> test in the module already does.
> Test-only change.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]