joseluisll opened a new pull request, #8787: URL: https://github.com/apache/hadoop/pull/8787
### Description of PR JIRA: https://issues.apache.org/jira/browse/YARN-12004 Supersedes #8786, which was closed when its branch was renamed. `TestRouterWebServicesREST`, and `TestFederationWebApp` which extends it, fail intermittently on CI because several tests don't wait for the RM's asynchronous application lifecycle. The change is test-only, in `TestRouterWebServicesREST.java`. - **`getAppAttempt`** read `getAttempts().get(0)` right after the submission returned, before the RM had created the first attempt, and threw `IndexOutOfBoundsException`. It now waits for the attempt, up to 10s. - **`testAppXML`, `testAppStateXML`, `testAppAttemptXML`, `testGetContainersXML`** each compared one Router read with one RM read of a value the RM was still changing (AM host, state, attempts, containers), e.g. `expected: <ACCEPTED> but was: <SUBMITTED>`. - A new **`assertRouterMatchesRM(path, type, key)`** reads the RM, then the Router, then the RM again, and compares a key extracted from each answer: - If the RM changed between its two reads, it retries. - If the RM was stable, the Router must match. A Router that disagrees with a stable RM twice fails at once, so a broken Router can't pass just because the RM's value eventually reaches whatever the Router reports. - **`testGetContainersXML`** retries stable-looking mismatches instead of failing fast. The endpoint reports only the current attempt's AM container, which lives 20–100 ms here, so it can appear and disappear within a single Router read, once per attempt. - A single-address `performGetCall` replaces the duplicated Router/RM request code, and every client it creates is closed. - The tests that now wait get a 30s timeout instead of 2s. ### How was this patch tested? - `TestRouterWebServicesREST` (42 tests) and `TestFederationWebApp` (52 tests) pass. - The four affected tests, looped 300 times in one JVM, failed none. - The comparison helper was exercised directly: a Router mismatching a stable RM fails after two tries (6 reads); an RM that keeps changing retries until the 10s timeout. ### For code changes: - [x] Does the title of this PR start with the corresponding JIRA issue id (e.g. 'HADOOP-17799. Your PR title ...')? - [ ] Object storage: N/A - [ ] If adding new dependencies to the code, are these dependencies licensed in a way that is compatible for inclusion under ASF 2.0? N/A, no new dependencies. - [ ] If applicable, have you updated the `LICENSE`, `LICENSE-binary`, `NOTICE-binary` files? N/A ### AI Tooling - [x] The PR includes the phrase "Contains content generated by Claude Code". - [x] My use of AI contributions follows the ASF legal policy https://www.apache.org/legal/generative-tooling.html Contains content generated by Claude Code. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
