[ 
https://issues.apache.org/jira/browse/DRILL-3808?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14877245#comment-14877245
 ] 

Aman Sinha commented on DRILL-3808:
-----------------------------------

Here's a better example: 
t2.tsv: (note: there is only 2 columns separated by \t on each row)
{code}
"Best Buy" Samsung 50" TV       2000.00
"Best Buy" Sharp 50" TV 1899.00
{code}

{code}
0: jdbc:drill:zk=local> select columns[0], columns[1] from `t2.tsv`;
Error: SYSTEM ERROR: TextParsingException: Error processing input: Cannot use 
newline character within quoted string, line=2, char=67. Content parsed: [ ]

Fragment 0:0

[Error Id: 26f21bb7-a7d5-4b31-8b12-9aec873fb278 on 192.168.1.105:31010]

  (com.univocity.parsers.common.TextParsingException) Error processing input: 
Cannot use newline character within quoted string, line=2, char=67. Content 
parsed: [ ]
    org.apache.drill.exec.store.easy.text.compliant.TextReader.parseNext():371
    
org.apache.drill.exec.store.easy.text.compliant.CompliantTextRecordReader.next():132
    org.apache.drill.exec.physical.impl.ScanBatch.next():183
{code}

> When reading TSV files, TextReader does not follow the standard
> ---------------------------------------------------------------
>
>                 Key: DRILL-3808
>                 URL: https://issues.apache.org/jira/browse/DRILL-3808
>             Project: Apache Drill
>          Issue Type: Bug
>          Components: Storage - Text & CSV
>            Reporter: Sean Hsuan-Yi Chu
>            Assignee: Sean Hsuan-Yi Chu
>            Priority: Critical
>
> According to references [1], [2]:
> In .csv, the double quote is a special character as it can optionally enclose 
> a text field. But in .tsv, it is not a special character, and it can appear 
> anywhere and when it does, it should treated as a literal. The tsv format 
> specification also does not provide for the tab or CR/LF characters to show 
> up anywhere in text fields. However, Drill treats tsv very the same like csv.
> For an example, given data:
> {code}
> "test"\t"test"
> {code}
> A query: select columns[0], columns[1] from `t.tsv`; Drill would give
> {code}
> test      test
> {code}
> However, according to the reference[2], it is supposed to be
> {code}
> "test"      "test"
> {code}
> Ideally, the Drill should follow the standard see[2].
> [1] CSV - https://tools.ietf.org/html/rfc4180
> [2] TSV - 
> http://www.iana.org/assignments/media-types/text/tab-separated-values



--
This message was sent by Atlassian JIRA
(v6.3.4#6332)

Reply via email to