[jira] [Updated] (MAPREDUCE-6558) multibyte delimiters with compressed input files generate duplicate records
[ https://issues.apache.org/jira/browse/MAPREDUCE-6558?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ] Jason Lowe updated MAPREDUCE-6558: -- Resolution: Fixed Hadoop Flags: Reviewed Fix Version/s: 2.6.5 2.7.3 2.8.0 Status: Resolved (was: Patch Available) Thanks, [~wilfreds]! I committed this to trunk, branch-2, branch-2.8, branch-2.7, and branch-2.6. > multibyte delimiters with compressed input files generate duplicate records > --- > > Key: MAPREDUCE-6558 > URL: https://issues.apache.org/jira/browse/MAPREDUCE-6558 > Project: Hadoop Map/Reduce > Issue Type: Bug > Components: mrv1, mrv2 >Affects Versions: 2.7.2 >Reporter: Wilfred Spiegelenburg >Assignee: Wilfred Spiegelenburg > Fix For: 2.8.0, 2.7.3, 2.6.5 > > Attachments: MAPREDUCE-6558.1.patch, MAPREDUCE-6558.2.patch, > MAPREDUCE-6558.3.patch > > > This is the follow up for MAPREDUCE-6549. Compressed files cause record > duplications as shown in different junit tests. The number of duplicated > records changes with the splitsize: > Unexpected number of records in split (splitsize = 10) > Expected: 41051 > Actual: 45062 > Unexpected number of records in split (splitsize = 10) > Expected: 41051 > Actual: 41052 > Test passes with splitsize = 147445 which is the compressed file length.The > file is a bzip2 file with 100k blocks and a total of 11 blocks -- This message was sent by Atlassian JIRA (v6.3.4#6332) - To unsubscribe, e-mail: mapreduce-issues-unsubscr...@hadoop.apache.org For additional commands, e-mail: mapreduce-issues-h...@hadoop.apache.org
[jira] [Updated] (MAPREDUCE-6558) multibyte delimiters with compressed input files generate duplicate records
[ https://issues.apache.org/jira/browse/MAPREDUCE-6558?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ] Wilfred Spiegelenburg updated MAPREDUCE-6558: - Attachment: MAPREDUCE-6558.3.patch > multibyte delimiters with compressed input files generate duplicate records > --- > > Key: MAPREDUCE-6558 > URL: https://issues.apache.org/jira/browse/MAPREDUCE-6558 > Project: Hadoop Map/Reduce > Issue Type: Bug > Components: mrv1, mrv2 >Affects Versions: 2.7.2 >Reporter: Wilfred Spiegelenburg >Assignee: Wilfred Spiegelenburg > Attachments: MAPREDUCE-6558.1.patch, MAPREDUCE-6558.2.patch, > MAPREDUCE-6558.3.patch > > > This is the follow up for MAPREDUCE-6549. Compressed files cause record > duplications as shown in different junit tests. The number of duplicated > records changes with the splitsize: > Unexpected number of records in split (splitsize = 10) > Expected: 41051 > Actual: 45062 > Unexpected number of records in split (splitsize = 10) > Expected: 41051 > Actual: 41052 > Test passes with splitsize = 147445 which is the compressed file length.The > file is a bzip2 file with 100k blocks and a total of 11 blocks -- This message was sent by Atlassian JIRA (v6.3.4#6332) - To unsubscribe, e-mail: mapreduce-issues-unsubscr...@hadoop.apache.org For additional commands, e-mail: mapreduce-issues-h...@hadoop.apache.org
[jira] [Updated] (MAPREDUCE-6558) multibyte delimiters with compressed input files generate duplicate records
[ https://issues.apache.org/jira/browse/MAPREDUCE-6558?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ] Wilfred Spiegelenburg updated MAPREDUCE-6558: - Attachment: MAPREDUCE-6558.2.patch > multibyte delimiters with compressed input files generate duplicate records > --- > > Key: MAPREDUCE-6558 > URL: https://issues.apache.org/jira/browse/MAPREDUCE-6558 > Project: Hadoop Map/Reduce > Issue Type: Bug > Components: mrv1, mrv2 >Affects Versions: 2.7.2 >Reporter: Wilfred Spiegelenburg >Assignee: Wilfred Spiegelenburg > Attachments: MAPREDUCE-6558.1.patch, MAPREDUCE-6558.2.patch > > > This is the follow up for MAPREDUCE-6549. Compressed files cause record > duplications as shown in different junit tests. The number of duplicated > records changes with the splitsize: > Unexpected number of records in split (splitsize = 10) > Expected: 41051 > Actual: 45062 > Unexpected number of records in split (splitsize = 10) > Expected: 41051 > Actual: 41052 > Test passes with splitsize = 147445 which is the compressed file length.The > file is a bzip2 file with 100k blocks and a total of 11 blocks -- This message was sent by Atlassian JIRA (v6.3.4#6332) - To unsubscribe, e-mail: mapreduce-issues-unsubscr...@hadoop.apache.org For additional commands, e-mail: mapreduce-issues-h...@hadoop.apache.org
[jira] [Updated] (MAPREDUCE-6558) multibyte delimiters with compressed input files generate duplicate records
[ https://issues.apache.org/jira/browse/MAPREDUCE-6558?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ] Wilfred Spiegelenburg updated MAPREDUCE-6558: - Attachment: MAPREDUCE-6558.1.patch A patch with test input that fails before the fix was made and passes after the fix was made. I have run all the tests that were there and they all still pass with this change. The tests have been run a large number of times with different input splits and none of them have failed. The use cases that passed before the fix was applied have been documented in the code. > multibyte delimiters with compressed input files generate duplicate records > --- > > Key: MAPREDUCE-6558 > URL: https://issues.apache.org/jira/browse/MAPREDUCE-6558 > Project: Hadoop Map/Reduce > Issue Type: Bug > Components: mrv1, mrv2 >Affects Versions: 2.7.2 >Reporter: Wilfred Spiegelenburg >Assignee: Wilfred Spiegelenburg > Attachments: MAPREDUCE-6558.1.patch > > > This is the follow up for MAPREDUCE-6549. Compressed files cause record > duplications as shown in different junit tests. The number of duplicated > records changes with the splitsize: > Unexpected number of records in split (splitsize = 10) > Expected: 41051 > Actual: 45062 > Unexpected number of records in split (splitsize = 10) > Expected: 41051 > Actual: 41052 > Test passes with splitsize = 147445 which is the compressed file length.The > file is a bzip2 file with 100k blocks and a total of 11 blocks -- This message was sent by Atlassian JIRA (v6.3.4#6332) - To unsubscribe, e-mail: mapreduce-issues-unsubscr...@hadoop.apache.org For additional commands, e-mail: mapreduce-issues-h...@hadoop.apache.org
[jira] [Updated] (MAPREDUCE-6558) multibyte delimiters with compressed input files generate duplicate records
[ https://issues.apache.org/jira/browse/MAPREDUCE-6558?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ] Wilfred Spiegelenburg updated MAPREDUCE-6558: - Status: Patch Available (was: Open) > multibyte delimiters with compressed input files generate duplicate records > --- > > Key: MAPREDUCE-6558 > URL: https://issues.apache.org/jira/browse/MAPREDUCE-6558 > Project: Hadoop Map/Reduce > Issue Type: Bug > Components: mrv1, mrv2 >Affects Versions: 2.7.2 >Reporter: Wilfred Spiegelenburg >Assignee: Wilfred Spiegelenburg > Attachments: MAPREDUCE-6558.1.patch > > > This is the follow up for MAPREDUCE-6549. Compressed files cause record > duplications as shown in different junit tests. The number of duplicated > records changes with the splitsize: > Unexpected number of records in split (splitsize = 10) > Expected: 41051 > Actual: 45062 > Unexpected number of records in split (splitsize = 10) > Expected: 41051 > Actual: 41052 > Test passes with splitsize = 147445 which is the compressed file length.The > file is a bzip2 file with 100k blocks and a total of 11 blocks -- This message was sent by Atlassian JIRA (v6.3.4#6332) - To unsubscribe, e-mail: mapreduce-issues-unsubscr...@hadoop.apache.org For additional commands, e-mail: mapreduce-issues-h...@hadoop.apache.org
[jira] [Updated] (MAPREDUCE-6558) multibyte delimiters with compressed input files generate duplicate records
[ https://issues.apache.org/jira/browse/MAPREDUCE-6558?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ] Junping Du updated MAPREDUCE-6558: -- Target Version/s: 2.8.0, 2.7.3, 2.6.5 (was: 2.8.0, 2.7.3, 2.6.4) > multibyte delimiters with compressed input files generate duplicate records > --- > > Key: MAPREDUCE-6558 > URL: https://issues.apache.org/jira/browse/MAPREDUCE-6558 > Project: Hadoop Map/Reduce > Issue Type: Bug > Components: mrv1, mrv2 >Affects Versions: 2.7.2 >Reporter: Wilfred Spiegelenburg >Assignee: Wilfred Spiegelenburg > > This is the follow up for MAPREDUCE-6549. Compressed files cause record > duplications as shown in different junit tests. The number of duplicated > records changes with the splitsize: > Unexpected number of records in split (splitsize = 10) > Expected: 41051 > Actual: 45062 > Unexpected number of records in split (splitsize = 10) > Expected: 41051 > Actual: 41052 > Test passes with splitsize = 147445 which is the compressed file length.The > file is a bzip2 file with 100k blocks and a total of 11 blocks -- This message was sent by Atlassian JIRA (v6.3.4#6332)
[jira] [Updated] (MAPREDUCE-6558) multibyte delimiters with compressed input files generate duplicate records
[ https://issues.apache.org/jira/browse/MAPREDUCE-6558?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ] Junping Du updated MAPREDUCE-6558: -- Target Version/s: 2.8.0, 2.7.3, 2.6.4 (was: 2.8.0, 2.6.3, 2.7.3) > multibyte delimiters with compressed input files generate duplicate records > --- > > Key: MAPREDUCE-6558 > URL: https://issues.apache.org/jira/browse/MAPREDUCE-6558 > Project: Hadoop Map/Reduce > Issue Type: Bug > Components: mrv1, mrv2 >Affects Versions: 2.7.2 >Reporter: Wilfred Spiegelenburg >Assignee: Wilfred Spiegelenburg > > This is the follow up for MAPREDUCE-6549. Compressed files cause record > duplications as shown in different junit tests. The number of duplicated > records changes with the splitsize: > Unexpected number of records in split (splitsize = 10) > Expected: 41051 > Actual: 45062 > Unexpected number of records in split (splitsize = 10) > Expected: 41051 > Actual: 41052 > Test passes with splitsize = 147445 which is the compressed file length.The > file is a bzip2 file with 100k blocks and a total of 11 blocks -- This message was sent by Atlassian JIRA (v6.3.4#6332)
[jira] [Updated] (MAPREDUCE-6558) multibyte delimiters with compressed input files generate duplicate records
[ https://issues.apache.org/jira/browse/MAPREDUCE-6558?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ] Tsuyoshi Ozawa updated MAPREDUCE-6558: -- Target Version/s: 2.8.0, 2.6.3, 2.7.3 > multibyte delimiters with compressed input files generate duplicate records > --- > > Key: MAPREDUCE-6558 > URL: https://issues.apache.org/jira/browse/MAPREDUCE-6558 > Project: Hadoop Map/Reduce > Issue Type: Bug > Components: mrv1, mrv2 >Affects Versions: 2.7.2 >Reporter: Wilfred Spiegelenburg >Assignee: Wilfred Spiegelenburg > > This is the follow up for MAPREDUCE-6549. Compressed files cause record > duplications as shown in different junit tests. The number of duplicated > records changes with the splitsize: > Unexpected number of records in split (splitsize = 10) > Expected: 41051 > Actual: 45062 > Unexpected number of records in split (splitsize = 10) > Expected: 41051 > Actual: 41052 > Test passes with splitsize = 147445 which is the compressed file length.The > file is a bzip2 file with 100k blocks and a total of 11 blocks -- This message was sent by Atlassian JIRA (v6.3.4#6332)