Md created TIKA-2593: ------------------------ Summary: docx with track change producing incorrect output Key: TIKA-2593 URL: https://issues.apache.org/jira/browse/TIKA-2593 Project: Tika Issue Type: Bug Components: core, handler Affects Versions: 1.17 Reporter: Md Attachments: sample.docx
I am using following code to extract text from docx file {code:java} contentHandler = new BodyContentHandler(); inputStream = new BufferedInputStream(new FileInputStream(inputFileName)); Metadata metadata = new Metadata(); StringBuilder fileContent = new StringBuilder(); recursiveParserWrapper.parse(inputStream, contentHandler, metadata, parseContext); System.out.println("Metadata WordCount Value: "+contentHandler.toString()); {code} When I am sending track revised files it's adding all the text deleted with the actual text and inserted text. Is there a way to tell parser to exclude the deleted text? Here is an example input Text: This is a sample text. -This part will- be deleted. +This is inserted.+ outputText: This is a sample text. This part will be deleted. This is inserted. Desired output: This is a sample text. be deleted. This is inserted. -- This message was sent by Atlassian JIRA (v7.6.3#76005)