[ 
https://issues.apache.org/jira/browse/HADOOP-19987?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18114998#comment-18114998
 ] 

ASF GitHub Bot commented on HADOOP-19987:
-----------------------------------------

joseluisll commented on PR #8724:
URL: https://github.com/apache/hadoop/pull/8724#issuecomment-5659736629

   @pan3793 @slfan1989 This PR is Green and ready to be reviewed. This one 
solves race conditions on minicluster when asked to be shutdown, it hanged on 
the closing datanodes. This surfaced because up to a month ago minicluster was 
not shutdown in the tests, leaked. I worked in the jiras to fix that. Properly 
closing the miniclusters and eliminating this race condition improves CI 
testing.
   
   It is a prerrequisite for 
[HADOOP-19979](https://issues.apache.org/jira/browse/HADOOP-19979), that will 
fix 4 flaky conditions so that we reduce later the GHA exclude list, as 
explained on [HADOOP-19981](https://issues.apache.org/jira/browse/HADOOP-19981) 
[umbrella for all GHA reactivation candidates from excluded-tests.txt] and its 
subtasks [specific identified candidates for reactivation].




> ObserverReadProxyProvider never falls back to the active NameNode when its 
> probing pool saturates
> -------------------------------------------------------------------------------------------------
>
>                 Key: HADOOP-19987
>                 URL: https://issues.apache.org/jira/browse/HADOOP-19987
>             Project: Hadoop Common
>          Issue Type: Sub-task
>          Components: ha, hdfs-client
>            Reporter: Jose Luis López
>            Priority: Minor
>              Labels: pull-request-available
>
> {{ObserverReadProxyProvider.getHAServiceStateWithTimeout}} catches
> {{RejectedExecutionException}} and returns null so the caller falls back to 
> the
> active NameNode:
> {code:java}
> } catch (RejectedExecutionException e) {
>   LOG.warn("Run out of threads to submit the request to query HA state. "
>       + "Ok to return null and we will fallback to use active NN to serve "
>       + "this request.");
>   return null;
> }
> {code}
> That fallback is unreachable.
> h3. Cause
> The pool is a {{BlockingThreadPoolExecutorService}}, which wraps the real 
> executor
> in a {{SemaphoredDelegatingExecutor}}. Its {{submit()}} blocks on
> {{queueingPermits.acquire()}} rather than rejecting, so it never throws
> {{RejectedExecutionException}} -- the handler above is dead code.
> Once the pool saturates, a caller that should have been shed onto the active 
> NN is
> instead parked. It parks while holding this provider's monitor, since the 
> probe is
> reached through the {{synchronized changeProxy()}}, so every other thread 
> needing
> that monitor waits behind it.
> h3. Fix
> Use a plain {{ThreadPoolExecutor}} with the same shape -- 4 threads, 128-deep 
> queue,
> 10s idle timeout via {{allowCoreThreadTimeOut}}, daemon threads -- and the 
> default
> {{AbortPolicy}}, so saturation reaches the fallback that was written for it. 
> The
> field type widens to {{ExecutorService}}.
> Behaviour is otherwise unchanged; the only observable difference is that 
> probing
> thread names gain a numeric suffix.
> h3. Verification
> Two tests added to {{TestObserverReadProxyProvider}}. Against the current
> {{BlockingThreadPoolExecutorService}} both fail, blocking for the full 20s
> deadline:
>   submit past capacity did not return within 20000ms: it blocked instead of 
> failing fast
>   HA state probe on a saturated pool did not return within 20000ms: it 
> blocked instead of failing fast
> With the fix the class passes 18/18, repeated three times.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to