Wilfred Spiegelenburg created MAPREDUCE-6558: ------------------------------------------------
Summary: multibyte delimiters with compressed input files generate duplicate records Key: MAPREDUCE-6558 URL: https://issues.apache.org/jira/browse/MAPREDUCE-6558 Project: Hadoop Map/Reduce Issue Type: Bug Components: mrv1, mrv2 Affects Versions: 2.7.2 Reporter: Wilfred Spiegelenburg Assignee: Wilfred Spiegelenburg This is the follow up for MAPREDUCE-6549. Compressed files cause record duplications as shown in different junit tests. The number of duplicated records changes with the splitsize: Unexpected number of records in split (splitsize = 10) Expected: 41051 Actual: 45062 Unexpected number of records in split (splitsize = 100000) Expected: 41051 Actual: 41052 Test passes with splitsize = 147445 which is the compressed file length.The file is a bzip2 file with 100k blocks and a total of 11 blocks -- This message was sent by Atlassian JIRA (v6.3.4#6332)