[
https://issues.apache.org/jira/browse/HADOOP-19982?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Jose Luis López updated HADOOP-19982:
-------------------------------------
Status: Patch Available (was: Open)
> Re-enable TestDFSClientRetries and TestBalancerWithHANameNodes in GHA
> ---------------------------------------------------------------------
>
> Key: HADOOP-19982
> URL: https://issues.apache.org/jira/browse/HADOOP-19982
> Project: Hadoop Common
> Issue Type: Sub-task
> Components: common, hdfs, test
> Reporter: Jose Luis López
> Assignee: Jose Luis López
> Priority: Major
> Labels: pull-request-available
>
> h3. What
> Remove two entries from \{{.github/gha-tests/exclude-tests.txt}}:
> * \{{**/org/apache/hadoop/hdfs/TestDFSClientRetries.java}}
> *
> \{{**/org/apache/hadoop/hdfs/server/balancer/TestBalancerWithHANameNodes.java}}
> h3. Why these two
> *TestDFSClientRetries* - fixed by HDFS-17972.
> {\{testLeaseRenewAndDFSOutputStreamDeadLock}} installed a
> {\{DFSClientFaultInjector}} whose \{{delayWhenRenewLeaseTimeout()}} blocked
> on an
> unbounded \{{testLatch.await()}}, and never restored the previous injector. If
> the test died before \{{NameNode.complete()}} counted the latch down, a lease
> renewer thread stayed parked inside the injector holding the \{{LeaseRenewer}}
> monitor. That hung the rest of the class until Surefire killed the fork,
> losing
> the class's results - exactly the "fails or aborts" the README describes. The
> fix bounds the wait at 30s and both counts the latch down and restores the
> injector in a \{{finally}} block. All 13 tests in the class already shut their
> clusters down in \{{finally}}.
> *TestBalancerWithHANameNodes* - fixed by HDFS-17957.
> {\{testBalancerWithHANameNodes}} and \{{testBalancerRequestSBNWithHA}} ran on
> a
> 60s budget when \{{waitStoragesNoStale}} alone can legitimately consume 60s,
> so
> the deadline could fire before the balancer even started on a loaded agent.
> Both are now 300s. The same change removed dead code that pinned the default
> NameNode RPC port.
> h3. Risk
> Both classes run in the \{{hdfs - other}} matrix entry
> (\{{-pl :hadoop-hdfs -DexcludedGroups=slow}}); neither carries
> \{{@Tag("slow")}},
> so one matrix entry exercises both removals. Neither class is excluded
> anywhere
> else in the repository. Since HDFS-17957 set
> {\{junit.jupiter.execution.timeout.default}} to 600s for this module, a
> regression in either class surfaces as a named test failure rather than a lost
> fork, so the cost of getting this wrong is one red test, not one red job.
> h3. Acceptance criteria
> Five consecutive successful GHA runs with both entries removed, per
> {\{.github/gha-tests/README.md}}.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]