[ 
https://issues.apache.org/jira/browse/TIKA-2347?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16264877#comment-16264877
 ] 

Hudson commented on TIKA-2347:
------------------------------

SUCCESS: Integrated in Jenkins build Tika-trunk #1396 (See 
[https://builds.apache.org/job/Tika-trunk/1396/])
Fix for TIKA-2347 Adds underline extraction from word documents (david: 
[https://github.com/apache/tika/commit/d64a32c63f376b9e003a4512adfec05414d4dfe6])
* (edit) 
tika-parsers/src/test/java/org/apache/tika/parser/microsoft/WordParserTest.java
* (edit) 
tika-parsers/src/test/java/org/apache/tika/parser/microsoft/ooxml/OOXMLParserTest.java
* (edit) 
tika-parsers/src/main/java/org/apache/tika/parser/microsoft/WordExtractor.java
* (edit) 
tika-parsers/src/main/java/org/apache/tika/parser/microsoft/ooxml/XWPFWordExtractorDecorator.java
TIKA-2347 - Added extraction of <strike> element in DOCX files (david: 
[https://github.com/apache/tika/commit/93cbed6df993ef01e59c55b86449b664e9052cae])
* (edit) 
tika-parsers/src/test/java/org/apache/tika/parser/microsoft/ooxml/OOXMLParserTest.java
* (edit) 
tika-parsers/src/main/java/org/apache/tika/parser/microsoft/ooxml/XWPFWordExtractorDecorator.java
* (edit) 
tika-parsers/src/test/java/org/apache/tika/parser/microsoft/WordParserTest.java
* (edit) tika-parsers/src/test/resources/test-documents/testWORD_various.docx
* (edit) tika-parsers/src/test/resources/test-documents/testWORD_various.doc
TIKA-2347 - Add underline extraction from Word documents (doc/docx) from 
(david: 
[https://github.com/apache/tika/commit/639f3bf361a08210da8fae68e3eeb4e12df6c4de])
* (edit) CHANGES.txt
TIKA-2347 - Add underline extraction from Word documents (doc/docx) from 
(david: 
[https://github.com/apache/tika/commit/beedc4277526c6327524acb3a799b2ca6c898a05])
* (edit) CHANGES.txt


> Underlined text is not decorated as such when extracting from word documents
> ----------------------------------------------------------------------------
>
>                 Key: TIKA-2347
>                 URL: https://issues.apache.org/jira/browse/TIKA-2347
>             Project: Tika
>          Issue Type: Bug
>          Components: parser
>    Affects Versions: 2.0, 1.14
>            Reporter: Stuart Hendren
>            Assignee: Dave Meikle
>             Fix For: 1.17
>
>
> When extracting from doc and docx bold and italic text decoration is 
> extracted, however underlining is not.  Can be demonstrated in WordParserTest 
> or OOXMLParserTest (change to docx) with the following test case.
> {code:title=WordParserTest.java|borderStyle=solid}
>     @Test
>     public void testTextDecoration() throws Exception {
>       XMLResult result = getXML("testWORD_various.doc");
>       String xml = result.xml;
>       assertTrue(xml.contains("<b>Bold</b>"));
>       assertTrue(xml.contains("<i>italic</i>"));
>       assertTrue(xml.contains("<u>underline</u>"));
>     }
> {code}



--
This message was sent by Atlassian JIRA
(v6.4.14#64029)

Reply via email to