[
https://issues.apache.org/jira/browse/MAPREDUCE-7539?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18108013#comment-18108013
]
ASF GitHub Bot commented on MAPREDUCE-7539:
-------------------------------------------
brumi1024 commented on code in PR #8556:
URL: https://github.com/apache/hadoop/pull/8556#discussion_r3855055879
##########
hadoop-yarn-project/hadoop-yarn/hadoop-yarn-common/src/main/java/org/apache/hadoop/yarn/logaggregation/filecontroller/ifile/LogAggregationIndexedFileController.java:
##########
@@ -237,12 +263,14 @@ public Object run() throws Exception {
} catch (Exception e) {
throw new IOException(e);
}
+ this.writeSession = session;
Review Comment:
I think this assignment happens too late to make initialization
exception-safe. If any filesystem operation throws after session.fsDataOStream
is opened, the new session is never assigned to writeSession.
AppLogAggregatorImpl will still call closeWriter() from its finally block, but
that method can only close the previous session or do nothing. This can leak
the output stream and its HDFS lease, leaving a partial file and potentially
blocking subsequent aggregation attempts.
Closing session.fsDataOStream directly when initialization fails, and
clearing writeSession in closeWriter() after cleanup could solve this.
> JobHistoryServer API fails to serve aggregated logs due to uuid mismatch
> ------------------------------------------------------------------------
>
> Key: MAPREDUCE-7539
> URL: https://issues.apache.org/jira/browse/MAPREDUCE-7539
> Project: Hadoop Map/Reduce
> Issue Type: Bug
> Components: jobhistoryserver
> Affects Versions: 3.5.0
> Reporter: Brian Goerlitz
> Assignee: Ferenc Erdelyi
> Priority: Major
> Labels: pull-request-available
>
> Due to Singleton annotations added in HADOOP-15984 for HSWebServices, the
> first time an ifile log is read via the {{/ws/v1/history/aggregatedlogs}}
> API, the UUID of the log is stored in the HSWebServices instance of the
> {{LogAggregationIndexedFileController}} and used for verification of all
> future log files. This results in failure to read any aggregated log files
> belonging to an app that is not the first one accessed after JHS restart.
> {noformat}
> 2026-06-08 20:16:40,368 WARN
> org.apache.hadoop.yarn.logaggregation.filecontroller.ifile.LogAggregationIndexedFileController:
> Can not get log meta from the log
> file:hdfs://nn:8020/tmp/logs/systest/bucket-logs-ifile/0002/application_1780935195539_0002/nm_8041
> The UUID from
> hdfs://nn:8020/tmp/logs/systest/bucket-logs-ifile/0002/application_1780935195539_0002/nm_8041
> is not correct. The offset of loaded UUID is 296605
> 2026-06-08 20:16:40,368 WARN
> org.apache.hadoop.yarn.webapp.GenericExceptionHandler: SERVICE_UNAVAILABLE
> javax.ws.rs.WebApplicationException: HTTP 500 Internal Server Error
> at
> org.apache.hadoop.yarn.server.webapp.LogServlet.getContainerLogMeta(LogServlet.java:134)
> at
> org.apache.hadoop.yarn.server.webapp.LogServlet.getContainerLogsInfo(LogServlet.java:325)
> at
> org.apache.hadoop.yarn.server.webapp.LogServlet.getLogsInfo(LogServlet.java:263)
> at
> org.apache.hadoop.mapreduce.v2.hs.webapp.HsWebServices.getAggregatedLogsMeta(HsWebServices.java:521)
> ...
> Caused by: org.apache.hadoop.yarn.webapp.NotFoundException: HTTP 404 Not Found
> at
> org.apache.hadoop.yarn.server.webapp.LogServlet.getContainerLogMeta(LogServlet.java:122)
> ... 86 more
> Caused by: java.lang.Exception: Can not get log meta for request.
> at
> org.apache.hadoop.yarn.webapp.NotFoundException.<init>(NotFoundException.java:45)
> ... 87 more
> {noformat}
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]