[
https://issues.apache.org/jira/browse/DRILL-3808?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14877167#comment-14877167
]
Jacques Nadeau commented on DRILL-3808:
---------------------------------------
I don't think there is a canonical version of either of these formats. There
may be in theory but I feel like people are pretty loose in interpretation.
Note the comments in univocity where they benchmark using the standard format
and beyond. [~amansinha100], the simpler reader definitely can't replace the
current reader as we've seen people need virtually all of the functionality of
the current reader. Also note that our reader is heavily built on the Univocity
stuff but it took substantial rework to make it work efficiently in our
framework. Building one more powerful reader instead of two separate ones
seemed more straightforward. We could always create a second reader but there
would be similar effort to get it to work efficiently within our memory
framework.
I generally think that more specific bugs should be opened rather than this
generic one. Then we can address the issue. These "standards" are not solid nor
well followed. Example specific issues could include:
> Someone wants to treats quotes as literals, this can be done by using an
> alternative quote character.
> Someone wants a faster reader: we should move to a word-wise reader instead
> of byte-wise. Forking into two separate readers may or may not be necessary
> to satisfy the performance criteria.
I think we'll be more productive if we focus on these kinds of specific issues.
What was the impetus behind filing this bug?
> When reading TSV files, TextReader does not follow the standard
> ---------------------------------------------------------------
>
> Key: DRILL-3808
> URL: https://issues.apache.org/jira/browse/DRILL-3808
> Project: Apache Drill
> Issue Type: Bug
> Components: Storage - Text & CSV
> Reporter: Sean Hsuan-Yi Chu
> Assignee: Sean Hsuan-Yi Chu
> Priority: Critical
>
> According to references [1], [2]:
> In .csv, the double quote is a special character as it can optionally enclose
> a text field. But in .tsv, it is not a special character, and it can appear
> anywhere and when it does, it should treated as a literal. The tsv format
> specification also does not provide for the tab or CR/LF characters to show
> up anywhere in text fields. However, Drill treats tsv very the same like csv.
> For an example, given data:
> {code}
> "test"\t"test"
> {code}
> A query: select columns[0], columns[1] from `t.tsv`; Drill would give
> {code}
> test test
> {code}
> However, according to the reference[2], it is supposed to be
> {code}
> "test" "test"
> {code}
> Ideally, the Drill should follow the standard see[2].
> [1] CSV - https://tools.ietf.org/html/rfc4180
> [2] TSV -
> http://www.iana.org/assignments/media-types/text/tab-separated-values
--
This message was sent by Atlassian JIRA
(v6.3.4#6332)