THere are two classes you can use to extract the text from a PDF document using PDFBox.
ExtractText is for extracting all the text, ExtractTextByArea will get text only in a certain area. The latter may be more effective if you are looking to strip out headers and footers on each page. Remember to write your code watching for page orientations. Duane Nickull *********************************** Technoracle Advanced Systems Inc. Consulting and Contracting; Proven Results! i. Neo4J, PDF, Java, LiveCycle ES, Flex, AIR, CQ5 & Mobile b. http://technoracle.blogspot.com t. @duanechaos "Don't fear the Graph! Embrace Neo4J" On 2012-11-12 8:25 AM, "aliki aliki" <[email protected]> wrote: > >Hello! >I would like to ask if there is any way that I can convert a pdf to plain >txt without including additional information of the pdf such as page >numbers, author names and titles written at the top or at the bottom of >each page. Thank you in advance!

