joseluisll opened a new pull request, #8785:
URL: https://github.com/apache/hadoop/pull/8785

   ### Description of PR
   
   JIRA: [YARN-12003](https://issues.apache.org/jira/browse/YARN-12003)
   
   `Router#serviceStart` schedules `SubClusterCleaner` on 
`scheduledExecutorService` at a fixed rate. This is on by default 
(`yarn.router.deregister.subcluster.enabled`) and runs every 60s. 
`Router#serviceStop` never shuts that executor down. As a result, the cleaner 
outlives the Router and keeps calling 
`FederationStateStoreFacade#getSubClusters` and `#deregisterSubCluster` on a 
state store that nothing maintains any more. Its thread is not a daemon thread, 
so it can also keep the JVM alive. This shows up when a Router is stopped 
inside a JVM that keeps running: tests, embedded or mock federation setups, or 
a stop/start in the same process.
   
   Changes:
   - `Router#serviceStop` now calls 
`HadoopExecutors#shutdown(scheduledExecutorService, LOG, 5, SECONDS)` before 
stopping the child services. This is a graceful shutdown with a bounded wait. 
It waits up to 5s for an in-flight cleaner run to finish, then calls 
`shutdownNow()` and waits up to 5s more, so stopping the Router takes at most 
10s longer. We don't call `shutdownNow()` straight away, so that a scan already 
in progress can finish before the state store it is reading is closed. 
`HadoopExecutors#shutdown` does nothing when the executor is null, so stopping 
a Router that was never initialised is still safe.
   - Adds a `@VisibleForTesting` getter, `Router#getScheduledExecutorService()`.
   - Adds the test `TestRouter#testServiceStopShutsDownScheduledExecutor`.
   
   ### How was this patch tested?
   
   - `TestRouter` on Linux with JDK 21: 6/6 tests pass, including the new 
`testServiceStopShutsDownScheduledExecutor`.
   - With the `HadoopExecutors.shutdown` call removed from `serviceStop`, the 
new test fails with `expected: <true> but was: <false>`. So the test does catch 
the missing shutdown.
   
   ### For code changes:
   
   - [x] Does the title of this PR start with the corresponding JIRA issue id 
(e.g. 'HADOOP-17799. Your PR title ...')?
   - [ ] Object storage: Have the integration tests been executed and the 
endpoint
         declared according to the connector-specific documentation? *Note: 
Automated CI
         testing doesn't cover all cases so manual testing with cloud storage 
is still
         required.*
   - [ ] If adding new dependencies to the code, are these dependencies 
licensed in a way that is compatible for inclusion under [ASF 
2.0](http://www.apache.org/legal/resolved.html#category-a)?
   - [ ] If applicable, have you updated the `LICENSE`, `LICENSE-binary`, 
`NOTICE-binary` files?
   
   ### AI Tooling
   
   If an AI tool was used:
   
   - [x] The PR includes the phrase "Contains content generated by <tool>"
         where <tool> is the name of the AI tool used.
   - [x] My use of AI contributions follows the ASF legal policy
         https://www.apache.org/legal/generative-tooling.html
   
   Contains content generated by Claude Code.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to