[ 
https://issues.apache.org/jira/browse/TIKA-2790?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16856099#comment-16856099
 ] 

Tim Allison commented on TIKA-2790:
-----------------------------------

[~kkrugler] I think I understand some of the string processing components that 
make it fast, but how does processing time not grow linearly if you aren't 
sampling?

||Detector||Length||Millis||Avg(ms)||Stdev||
|YalderDetector|10|4442|0.06|1.19|
|YalderDetector|50|13017|0.17|0.4|
|YalderDetector|100|14149|0.19|0.41|
|YalderDetector|200|14686|0.2|0.41|
|YalderDetector|500|14536|0.19|0.43|
|YalderDetector|1000|14993|0.2|0.41|
|YalderDetector|5000|16627|0.22|0.43|
|YalderDetector|10000|18884|0.25|0.46|
|YalderDetector|20000|20702|0.28|0.48|
|YalderDetector|50000|21749|0.29|0.48|
|YalderDetector|100000|23503|0.32|0.49|

 

> Consider switching lang-detection in tika-eval to open-nlp
> ----------------------------------------------------------
>
>                 Key: TIKA-2790
>                 URL: https://issues.apache.org/jira/browse/TIKA-2790
>             Project: Tika
>          Issue Type: Improvement
>            Reporter: Tim Allison
>            Priority: Major
>         Attachments: fra_mixed_100000_0.0_0.txt, langid_20190509.zip, 
> langid_20190510.zip, langid_20190514.zip, langid_20190514_plus_minus_1.zip
>
>




--
This message was sent by Atlassian JIRA
(v7.6.3#76005)

Reply via email to