[ 
https://issues.apache.org/jira/browse/TIKA-1300?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14045881#comment-14045881
 ] 

Tim Allison commented on TIKA-1300:
-----------------------------------

[~tilman], [~tboehme] and [~msahyoun], thank you all for your feedback.

The insights above will be helpful to users choosing which parser to use as 
default.

 [~msahyoun], thank you for actually opening the documents and performing 
manual review.  That offers some very useful insight.

My takeaway is that I should probably re-open TIKA-1205 and implement that at 
some point.

All, if this exercise was of use/interest to you, I'd like to invite you and 
all of your PDFBox colleagues to participate in TIKA-1302.  That is still very 
nascent, and all contributions will be welcome. Initially, recommendations on 
corpora, metrics and methods would be handy.  Longer term, if you wanted to 
borrow code and/or resources to standup your own ongoing batch eval system or 
collaborate with us on ours, I think that would be a mutual benefit for both of 
our projects.

Thank you, again!

> Switch default PDFBox parser to NonSequentialParser
> ---------------------------------------------------
>
>                 Key: TIKA-1300
>                 URL: https://issues.apache.org/jira/browse/TIKA-1300
>             Project: Tika
>          Issue Type: Improvement
>          Components: parser
>            Reporter: Tim Allison
>            Assignee: Tim Allison
>            Priority: Minor
>             Fix For: 1.7
>
>         Attachments: tika_1_6_ClassicsVsNonSeq.zip
>
>
> On TIKA-1298, [~tilman] recommended switching Tika's default to the 
> NonSequentialParser. We added a parameter to use the NonSequentialParser in 
> TIKA-1201, and there's some good discussion there about the benefits.
> Is the community in favor of switching the default now?



--
This message was sent by Atlassian JIRA
(v6.2#6252)

Reply via email to