Md created TIKA-2593:
------------------------

             Summary: docx with track change producing incorrect output
                 Key: TIKA-2593
                 URL: https://issues.apache.org/jira/browse/TIKA-2593
             Project: Tika
          Issue Type: Bug
          Components: core, handler
    Affects Versions: 1.17
            Reporter: Md
         Attachments: sample.docx

I am using following code to extract text from docx file 
{code:java}
contentHandler = new BodyContentHandler();
inputStream = new BufferedInputStream(new FileInputStream(inputFileName));
Metadata metadata = new Metadata();
StringBuilder fileContent = new StringBuilder();
recursiveParserWrapper.parse(inputStream, contentHandler, metadata, 
parseContext);
System.out.println("Metadata WordCount Value: "+contentHandler.toString());

{code}
When I am sending track revised files it's adding all the text deleted with the 
actual text and inserted text. Is there a way to tell parser to exclude the 
deleted text?

Here is an example 

input Text: This is a sample text. -This part will- be deleted. +This is 
inserted.+

outputText: This is a sample text. This part will be deleted. This is inserted.

Desired output: This is a sample text.  be deleted. This is inserted.



--
This message was sent by Atlassian JIRA
(v7.6.3#76005)

Reply via email to