[ https://issues.apache.org/jira/browse/YARN-2001?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14137644#comment-14137644 ]
Vinod Kumar Vavilapalli commented on YARN-2001: ----------------------------------------------- This looks almost close except for the logging - we don't have any indication of this wait in the RM logs. > Threshold for RM to accept requests from AM after failover > ---------------------------------------------------------- > > Key: YARN-2001 > URL: https://issues.apache.org/jira/browse/YARN-2001 > Project: Hadoop YARN > Issue Type: Sub-task > Components: resourcemanager > Reporter: Jian He > Assignee: Jian He > Attachments: YARN-2001.1.patch, YARN-2001.2.patch, YARN-2001.3.patch, > YARN-2001.4.patch > > > After failover, RM may require a certain threshold to determine whether it’s > safe to make scheduling decisions and start accepting new container requests > from AMs. The threshold could be a certain amount of nodes. i.e. RM waits > until a certain amount of nodes joining before accepting new container > requests. Or it could simply be a timeout, only after the timeout RM accepts > new requests. > NMs joined after the threshold can be treated as new NMs and instructed to > kill all its containers. -- This message was sent by Atlassian JIRA (v6.3.4#6332)