[jira] [Updated] (TIKA-2038) A more accurate facility for detecting Charset Encoding of HTML documents

Ken Krugler (JIRA) Thu, 21 Jul 2016 15:40:29 -0700

     [ 
https://issues.apache.org/jira/browse/TIKA-2038?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]


Ken Krugler updated TIKA-2038:
------------------------------
    Description: 
Currently, Tika uses icu4j for detecting charset encoding of HTML documents as 
well as the other naturally text documents. But the accuracy of encoding 
detector tools, including icu4j, in dealing with the HTML documents is 
meaningfully less than from which the other text documents. Hence, in our 
project I developed a library that works pretty well for HTML documents, which 
is available here: https://github.com/shabanali-faghani/IUST-HTMLCharDet

Since Tika is widely used with and within some of other Apache stuffs such as 
Nutch, Lucene, Solr, etc. and these projects are strongly in connection with 
the HTML documents, it seems that having such an facility in Tika also will 
help them to become more accurate.

  was:
Currently, Tika uses icu4j for detecting charset encoding of HTML documents as 
well as the other naturally text documents. But the accuracy of encoding 
detector tools, including icu4j, in dealing with the HTML documents is 
meaningfully less than from which the other text documents. Hence, in our 
project I developed a library that works pretty well for HTML documents, which 
is available here: https://github.com/shabanali-faghani/IUST-HTMLCharDet
Since Tika is widely used with and within some of other Apache stuffs such as 
Nutch, Lucene, Solr, etc. and these projects are strongly in connection with 
the HTML documents, it seems that having such an facility in Tika also will 
help them to become more accurate.

     Issue Type: Improvement  (was: New Feature)

> A more accurate facility for detecting Charset Encoding of HTML documents
> -------------------------------------------------------------------------
>
>                 Key: TIKA-2038
>                 URL: https://issues.apache.org/jira/browse/TIKA-2038
>             Project: Tika
>          Issue Type: Improvement
>          Components: core, detector
>            Reporter: Shabanali Faghani
>            Priority: Minor
>
> Currently, Tika uses icu4j for detecting charset encoding of HTML documents 
> as well as the other naturally text documents. But the accuracy of encoding 
> detector tools, including icu4j, in dealing with the HTML documents is 
> meaningfully less than from which the other text documents. Hence, in our 
> project I developed a library that works pretty well for HTML documents, 
> which is available here: https://github.com/shabanali-faghani/IUST-HTMLCharDet
> Since Tika is widely used with and within some of other Apache stuffs such as 
> Nutch, Lucene, Solr, etc. and these projects are strongly in connection with 
> the HTML documents, it seems that having such an facility in Tika also will 
> help them to become more accurate.



--
This message was sent by Atlassian JIRA
(v6.3.4#6332)

[jira] [Updated] (TIKA-2038) A more accurate facility for detecting Charset Encoding of HTML documents

Reply via email to