[ https://issues.apache.org/jira/browse/MAPREDUCE-5308?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=13677601#comment-13677601 ]
Hadoop QA commented on MAPREDUCE-5308: -------------------------------------- {color:green}+1 overall{color}. Here are the results of testing the latest attachment http://issues.apache.org/jira/secure/attachment/12586598/MAPREDUCE-5308.patch against trunk revision . {color:green}+1 @author{color}. The patch does not contain any @author tags. {color:green}+1 tests included{color}. The patch appears to include 1 new or modified test files. {color:green}+1 javac{color}. The applied patch does not increase the total number of javac compiler warnings. {color:green}+1 javadoc{color}. The javadoc tool did not generate any warning messages. {color:green}+1 eclipse:eclipse{color}. The patch built with eclipse:eclipse. {color:green}+1 findbugs{color}. The patch does not introduce any new Findbugs (version 1.3.9) warnings. {color:green}+1 release audit{color}. The applied patch does not increase the total number of release audit warnings. {color:green}+1 core tests{color}. The patch passed unit tests in hadoop-mapreduce-project/hadoop-mapreduce-client/hadoop-mapreduce-client-core. {color:green}+1 contrib tests{color}. The patch passed contrib unit tests. Test results: https://builds.apache.org/job/PreCommit-MAPREDUCE-Build/3746//testReport/ Console output: https://builds.apache.org/job/PreCommit-MAPREDUCE-Build/3746//console This message is automatically generated. > Shuffling to memory can get out-of-sync when fetching multiple compressed map > outputs > ------------------------------------------------------------------------------------- > > Key: MAPREDUCE-5308 > URL: https://issues.apache.org/jira/browse/MAPREDUCE-5308 > Project: Hadoop Map/Reduce > Issue Type: Bug > Affects Versions: trunk, 2.0.3-alpha, 0.23.8 > Reporter: Nathan Roberts > Assignee: Nathan Roberts > Attachments: MAPREDUCE-5308.patch > > > When a reducer is fetching multiple compressed map outputs from a host, the > fetcher can get out-of-sync with the IFileInputStream, causing several of the > maps to fail to fetch. > This occurs because decompressors can return all the decompressed bytes > before actually processing all the bytes in the compressed stream (due to > checksums or other trailing data that we ignore). In the unfortunate case > where these extra bytes cross an io.file.buffer.size boundary, some extra > bytes will be left over and the next map_output will not fetch correctly > (usually due to an invalid map_id). > This scenario is not typically fatal to a job because the failure is charged > to the map_output immediately following the "bad" one and the subsequent > retry will normally work. -- This message is automatically generated by JIRA. If you think it was sent incorrectly, please contact your JIRA administrators For more information on JIRA, see: http://www.atlassian.com/software/jira