gerashegalov opened a new pull request, #8770:
URL: https://github.com/apache/hadoop/pull/8770

   ### Description of PR
   
   Fixes [YARN-11995](https://issues.apache.org/jira/browse/YARN-11995).
   
   Container history currently drops custom allocations such as `yarn.io/gpu`: 
the timeline publishers store only memory and vcores, and the history readers 
reconstruct a two-resource report. An allocation of two GPUs therefore appears 
as zero in history.
   
   Persist custom allocation values and units in the additive 
`YARN_CONTAINER_ALLOCATED_RESOURCES` info field. Update both ResourceManager 
timeline publishers, the NodeManager v2 publisher, and both history report 
converters. Shared helpers snapshot allocations before asynchronous 
publication, accept integer/long JSON values, and convert stored units to the 
reader's configured units.
   
   Existing memory/vcore fields and legacy entities remain compatible. Unknown 
resource types are skipped with a warning while their raw timeline data remains 
available. Document the resource-type configuration required for history 
servers and standalone clients. Allocations omitted by older publishers cannot 
be recovered retroactively.
   
   ### How was this patch tested?
   
   - 107 tests passed on JDK 17, with zero failures, errors, or skips, across 
`TestTimelineServiceHelper`, `TestSystemMetricsPublisher`, 
`TestSystemMetricsPublisherForV2`, `TestNMTimelinePublisher`, 
`TestApplicationHistoryManagerOnTimelineStore`, `TestAHSWebServices`, and 
`TestAHSv2ClientImpl`. Maven ran with `-Dmaven.test.failure.ignore=false`.
   - Coverage includes JSON/store round trips, single-container and list REST 
responses, GPU and other custom resources, zero and large values, unit 
conversion, legacy entities, unknown resource types, and source-mutation 
isolation.
   - Additional standalone probes passed for the generic storage codec, 
creation/finish entity merging, container-report protobuf serialization, and 
fresh-JVM loading of `resource-types.xml`.
   - `git diff --check` passed. Checkstyle reported 45 existing diagnostics in 
the checked files, none on changed lines.
   - A live HBase-backed ATSv2 cluster test was not run.
   
   ### For code changes:
   
   - [x] Does the title of this PR start with the corresponding JIRA issue id 
(e.g. 'HADOOP-17799. Your PR title ...')?
   - [ ] Object storage: Have the integration tests been executed and the 
endpoint
         declared according to the connector-specific documentation? *Note: 
Automated CI
         testing doesn't cover all cases so manual testing with cloud storage 
is still
         required.* Not applicable: no object storage changes.
   - [ ] If adding new dependencies to the code, are these dependencies 
licensed in a way that is compatible for inclusion under [ASF 
2.0](http://www.apache.org/legal/resolved.html#category-a)? Not applicable: no 
new dependencies.
   - [ ] If applicable, have you updated the `LICENSE`, `LICENSE-binary`, 
`NOTICE-binary` files? Not applicable: no licensing changes.
   
   ### AI Tooling
   
   Contains content generated by OpenAI Codex.
   
   If an AI tool was used:
   
   - [x] The PR includes the phrase "Contains content generated by <tool>"
         where <tool> is the name of the AI tool used.
   - [x] My use of AI contributions follows the ASF legal policy
         https://www.apache.org/legal/generative-tooling.html
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to