[ 
https://issues.apache.org/jira/browse/HADOOP-19987?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Jose Luis López reassigned HADOOP-19987:
----------------------------------------

    Assignee: Jose Luis López

> ObserverReadProxyProvider never falls back to the active NameNode when its 
> probing pool saturates
> -------------------------------------------------------------------------------------------------
>
>                 Key: HADOOP-19987
>                 URL: https://issues.apache.org/jira/browse/HADOOP-19987
>             Project: Hadoop Common
>          Issue Type: Sub-task
>          Components: ha, hdfs-client
>            Reporter: Jose Luis López
>            Assignee: Jose Luis López
>            Priority: Minor
>              Labels: pull-request-available
>
> {{ObserverReadProxyProvider.getHAServiceStateWithTimeout}} catches
> {{RejectedExecutionException}} and returns null so the caller falls back to 
> the
> active NameNode:
> {code:java}
> } catch (RejectedExecutionException e) {
>   LOG.warn("Run out of threads to submit the request to query HA state. "
>       + "Ok to return null and we will fallback to use active NN to serve "
>       + "this request.");
>   return null;
> }
> {code}
> That fallback is unreachable.
> h3. Cause
> The pool is a {{BlockingThreadPoolExecutorService}}, which wraps the real 
> executor
> in a {{SemaphoredDelegatingExecutor}}. Its {{submit()}} blocks on
> {{queueingPermits.acquire()}} rather than rejecting, so it never throws
> {{RejectedExecutionException}} -- the handler above is dead code.
> Once the pool saturates, a caller that should have been shed onto the active 
> NN is
> instead parked. It parks while holding this provider's monitor, since the 
> probe is
> reached through the {{synchronized changeProxy()}}, so every other thread 
> needing
> that monitor waits behind it.
> h3. Fix
> Use a plain {{ThreadPoolExecutor}} with the same shape -- 4 threads, 128-deep 
> queue,
> 10s idle timeout via {{allowCoreThreadTimeOut}}, daemon threads -- and the 
> default
> {{AbortPolicy}}, so saturation reaches the fallback that was written for it. 
> The
> field type widens to {{ExecutorService}}.
> Behaviour is otherwise unchanged; the only observable difference is that 
> probing
> thread names gain a numeric suffix.
> h3. Verification
> Two tests added to {{TestObserverReadProxyProvider}}. Against the current
> {{BlockingThreadPoolExecutorService}} both fail, blocking for the full 20s
> deadline:
>   submit past capacity did not return within 20000ms: it blocked instead of 
> failing fast
>   HA state probe on a saturated pool did not return within 20000ms: it 
> blocked instead of failing fast
> With the fix the class passes 18/18, repeated three times.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to