[ 
https://issues.apache.org/jira/browse/PDFBOX-588?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=12981910#action_12981910
 ] 

Hesham commented on PDFBOX-588:
-------------------------------

I do not know what is a fragmented font !
But i have created a sample project to test extracting text from the PDF 
reference, and it took the same time i mentioned for the 2 PDFBox versions. I 
do not understand how it works fine with you !

Here is my code :
private void readPDFButtonActionPerformed() {
        try {
                PDDocument pdfRef = PDDocument.load( 
"C:\\pdf_reference_1.7.pdf" );
                PDFTextStripper stripper = new PDFTextStripper();
                
                for( int pageNum = 1; pageNum < pdfRef.getNumberOfPages(); 
pageNum++ ) {
                        System.out.println( pageNum );
                        stripper.setStartPage( pageNum );
                        stripper.setEndPage( pageNum );
                        stripper.getText( pdfRef ); 
                }
                System.out.println( "Done" );
        } catch (IOException e) {
                e.printStackTrace();
        }
}

> Problem extracting text in newline characters
> ---------------------------------------------
>
>                 Key: PDFBOX-588
>                 URL: https://issues.apache.org/jira/browse/PDFBOX-588
>             Project: PDFBox
>          Issue Type: Bug
>          Components: Text extraction
>    Affects Versions: 0.8.0-incubator, 1.3.1, 1.4.0
>         Environment: Win XP
>            Reporter: Hesham
>            Assignee: Andreas Lehmkühler
>         Attachments: Enters-sample.pdf, PDFBOX588-Enters-sample.txt, 
> PDFBOX588-Enters-sample1.png, PDFBOX588-Enters-sample1.png, 
> PDFTextStripper.patch
>
>
> Hello ,
>  
> I have a PDF file with 1 page only, when I try to extract its text using :
> String pageData = stripper.getText( pdfFile );
> It ignores some Enter characters between lines, so the last word in the line 
> and the first word in the next line appear as 1 word without spaces between 
> them !!
> While if I copy the PDF text manually from the PDF and paste it in a text 
> editor, Enter characters appear after the same lines that caused the problem 
> in PDFBox.
> Please check the attached file as a sample.
>  
> Is there a way to fix this ?
>  
> Best regards ,

-- 
This message is automatically generated by JIRA.
-
You can reply to this email to add a comment to the issue online.

Reply via email to