[ https://issues.apache.org/jira/browse/TIKA-2100?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16490959#comment-16490959 ]
ASF GitHub Bot commented on TIKA-2100: -------------------------------------- GerardBouchar commented on issue #238: TIKA-2100 extract content language from html lang attribute URL: https://github.com/apache/tika/pull/238#issuecomment-392110416 I reverted the automatic changes to the import statements (there are unused imports, though). But I don't see how to add the language information to the output HTML without fiddling with XHTMLContentHandler... ---------------------------------------------------------------- This is an automated message from the Apache Git Service. To respond to the message, please log on GitHub and use the URL above to go to the specific comment. For queries about this service, please contact Infrastructure at: us...@infra.apache.org > Html Parser does not keep the html tag attributes > ------------------------------------------------- > > Key: TIKA-2100 > URL: https://issues.apache.org/jira/browse/TIKA-2100 > Project: Tika > Issue Type: Bug > Components: parser > Affects Versions: 1.13 > Reporter: Gerard Bouchar > Priority: Major > > Parsing a very simple html like > <!DOCTYPE html> > <html lang="en"> > <head> > <title>Page Title</title> > </head> > <body> > <h1 align="left">My First Heading</h1> > <p>My first paragraph.</p> > </body> > </html> > you won't be able to access the html tag's attributes (here lang="en") in the > ContentHandler : > *in the method startElement(String ns, String localName, String name, > Attributes atts), atts is empty. > *Moreover it seems that the html tag's attributes are not passed trough the > HtmlMapper.mapSafeAttribute method too. -- This message was sent by Atlassian JIRA (v7.6.3#76005)