[
https://issues.apache.org/jira/browse/HADOOP-19987?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Jose Luis López reassigned HADOOP-19987:
----------------------------------------
Assignee: Jose Luis López
> ObserverReadProxyProvider never falls back to the active NameNode when its
> probing pool saturates
> -------------------------------------------------------------------------------------------------
>
> Key: HADOOP-19987
> URL: https://issues.apache.org/jira/browse/HADOOP-19987
> Project: Hadoop Common
> Issue Type: Sub-task
> Components: ha, hdfs-client
> Reporter: Jose Luis López
> Assignee: Jose Luis López
> Priority: Minor
> Labels: pull-request-available
>
> {{ObserverReadProxyProvider.getHAServiceStateWithTimeout}} catches
> {{RejectedExecutionException}} and returns null so the caller falls back to
> the
> active NameNode:
> {code:java}
> } catch (RejectedExecutionException e) {
> LOG.warn("Run out of threads to submit the request to query HA state. "
> + "Ok to return null and we will fallback to use active NN to serve "
> + "this request.");
> return null;
> }
> {code}
> That fallback is unreachable.
> h3. Cause
> The pool is a {{BlockingThreadPoolExecutorService}}, which wraps the real
> executor
> in a {{SemaphoredDelegatingExecutor}}. Its {{submit()}} blocks on
> {{queueingPermits.acquire()}} rather than rejecting, so it never throws
> {{RejectedExecutionException}} -- the handler above is dead code.
> Once the pool saturates, a caller that should have been shed onto the active
> NN is
> instead parked. It parks while holding this provider's monitor, since the
> probe is
> reached through the {{synchronized changeProxy()}}, so every other thread
> needing
> that monitor waits behind it.
> h3. Fix
> Use a plain {{ThreadPoolExecutor}} with the same shape -- 4 threads, 128-deep
> queue,
> 10s idle timeout via {{allowCoreThreadTimeOut}}, daemon threads -- and the
> default
> {{AbortPolicy}}, so saturation reaches the fallback that was written for it.
> The
> field type widens to {{ExecutorService}}.
> Behaviour is otherwise unchanged; the only observable difference is that
> probing
> thread names gain a numeric suffix.
> h3. Verification
> Two tests added to {{TestObserverReadProxyProvider}}. Against the current
> {{BlockingThreadPoolExecutorService}} both fail, blocking for the full 20s
> deadline:
> submit past capacity did not return within 20000ms: it blocked instead of
> failing fast
> HA state probe on a saturated pool did not return within 20000ms: it
> blocked instead of failing fast
> With the fix the class passes 18/18, repeated three times.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]