[
https://issues.apache.org/jira/browse/PDFBOX-1769?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=13842497#comment-13842497
]
Andreas Lehmkühler commented on PDFBOX-1769:
--------------------------------------------
I added a fix for the stream issue in revisions 1549025 and 1549027. The
non-sequential parser now detects a wrong stream length and uses the "old"
approach readUntilEndStream to read the data.
Now the pdf opens, but has some other issues.
> Fix crash on invalid xref
> -------------------------
>
> Key: PDFBOX-1769
> URL: https://issues.apache.org/jira/browse/PDFBOX-1769
> Project: PDFBox
> Issue Type: Wish
> Components: Parsing
> Affects Versions: 1.8.2
> Reporter: William Palmer
> Assignee: Andreas Lehmkühler
>
> Need to search for a correct xref start address
> Example file:
> http://digitalcorpora.org/corp/nps/files/govdocs1/020/020747.pdf
> Exception in thread "main" java.io.IOException: Error: Expected an integer
> type, actual='ref'
> at org.apache.pdfbox.pdfparser.BaseParser.readInt(BaseParser.java:1622)
> Using the code:
> PDFTextStripper ts = new PDFTextStripper();
> PrintWriter out = new PrintWriter(new FileWriter(new File (pFile+".txt")));
> RandomAccess scratchFile = new
> RandomAccessFile(File.createTempFile("pdfbox-", ".tmp"), "rw");
> PDDocument doc = PDDocument.loadNonSeq(new File(pFile), scratchFile)
> ts.setForceParsing(true);
> ts.writeText(doc, out);
> Related: PDFBOX-1757
--
This message was sent by Atlassian JIRA
(v6.1#6144)