[jira] Commented: (HADOOP-3144) better fault tolerance for corrupted text files

dhruba borthakur (JIRA) Sun, 06 Apr 2008 23:25:07 -0700

    [ 
https://issues.apache.org/jira/browse/HADOOP-3144?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=12586249#action_12586249
 ]


dhruba borthakur commented on HADOOP-3144:
------------------------------------------

Hi Zheng, I see that there are many changes in the patch that are changes only 
in whitespace characters. This makes the patch longer and is harder to review. 
Will it be possible for you to re-submit this patch with whitespace-only 
changes to lines removed? thanks.

> better fault tolerance for corrupted text files
> -----------------------------------------------
>
>                 Key: HADOOP-3144
>                 URL: https://issues.apache.org/jira/browse/HADOOP-3144
>             Project: Hadoop Core
>          Issue Type: Bug
>          Components: mapred
>    Affects Versions: 0.15.3
>            Reporter: Joydeep Sen Sarma
>            Assignee: Zheng Shao
>         Attachments: 3144.patch
>
>
> every once in a while - we encounter corrupted text files (corrupted at 
> source prior to copying into hadoop). inevitably - some of the data looks 
> like a really really long line and hadoop trips over trying to stuff it into 
> an in memory object and gets outofmem error. Code looks same way in trunk as 
> well .. 
> so looking for an option to the textinputformat (and like) to ignore long 
> lines. ideally - we would just skip errant lines above a certain size limit.

-- 
This message is automatically generated by JIRA.
-
You can reply to this email to add a comment to the issue online.

[jira] Commented: (HADOOP-3144) better fault tolerance for corrupted text files

Reply via email to