[ https://issues.apache.org/jira/browse/TIKA-2347?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16264877#comment-16264877 ]
Hudson commented on TIKA-2347: ------------------------------ SUCCESS: Integrated in Jenkins build Tika-trunk #1396 (See [https://builds.apache.org/job/Tika-trunk/1396/]) Fix for TIKA-2347 Adds underline extraction from word documents (david: [https://github.com/apache/tika/commit/d64a32c63f376b9e003a4512adfec05414d4dfe6]) * (edit) tika-parsers/src/test/java/org/apache/tika/parser/microsoft/WordParserTest.java * (edit) tika-parsers/src/test/java/org/apache/tika/parser/microsoft/ooxml/OOXMLParserTest.java * (edit) tika-parsers/src/main/java/org/apache/tika/parser/microsoft/WordExtractor.java * (edit) tika-parsers/src/main/java/org/apache/tika/parser/microsoft/ooxml/XWPFWordExtractorDecorator.java TIKA-2347 - Added extraction of <strike> element in DOCX files (david: [https://github.com/apache/tika/commit/93cbed6df993ef01e59c55b86449b664e9052cae]) * (edit) tika-parsers/src/test/java/org/apache/tika/parser/microsoft/ooxml/OOXMLParserTest.java * (edit) tika-parsers/src/main/java/org/apache/tika/parser/microsoft/ooxml/XWPFWordExtractorDecorator.java * (edit) tika-parsers/src/test/java/org/apache/tika/parser/microsoft/WordParserTest.java * (edit) tika-parsers/src/test/resources/test-documents/testWORD_various.docx * (edit) tika-parsers/src/test/resources/test-documents/testWORD_various.doc TIKA-2347 - Add underline extraction from Word documents (doc/docx) from (david: [https://github.com/apache/tika/commit/639f3bf361a08210da8fae68e3eeb4e12df6c4de]) * (edit) CHANGES.txt TIKA-2347 - Add underline extraction from Word documents (doc/docx) from (david: [https://github.com/apache/tika/commit/beedc4277526c6327524acb3a799b2ca6c898a05]) * (edit) CHANGES.txt > Underlined text is not decorated as such when extracting from word documents > ---------------------------------------------------------------------------- > > Key: TIKA-2347 > URL: https://issues.apache.org/jira/browse/TIKA-2347 > Project: Tika > Issue Type: Bug > Components: parser > Affects Versions: 2.0, 1.14 > Reporter: Stuart Hendren > Assignee: Dave Meikle > Fix For: 1.17 > > > When extracting from doc and docx bold and italic text decoration is > extracted, however underlining is not. Can be demonstrated in WordParserTest > or OOXMLParserTest (change to docx) with the following test case. > {code:title=WordParserTest.java|borderStyle=solid} > @Test > public void testTextDecoration() throws Exception { > XMLResult result = getXML("testWORD_various.doc"); > String xml = result.xml; > assertTrue(xml.contains("<b>Bold</b>")); > assertTrue(xml.contains("<i>italic</i>")); > assertTrue(xml.contains("<u>underline</u>")); > } > {code} -- This message was sent by Atlassian JIRA (v6.4.14#64029)