[ 
https://issues.apache.org/jira/browse/TIKA-4891?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18114677#comment-18114677
 ] 

Tim Allison commented on TIKA-4891:
-----------------------------------

Sorry, y, I was not being precise. PDF/UA is more specific. The point of this 
ticket is to do a better job with extracting structural markup into our output 
as long as the tags are not pathological.

Tika had a separate handler that I wrote a while ago before I understood well 
how the tags work and how to integrate that information with our+PDFBox's 
current extraction algorithm.

The point of this ticket is to figure out a better way to extract that 
structure when it exists – with guards (as possible) against tags that are nutz.

> Improve handling of PDF/UA structural tags/marked content
> ---------------------------------------------------------
>
>                 Key: TIKA-4891
>                 URL: https://issues.apache.org/jira/browse/TIKA-4891
>             Project: Tika
>          Issue Type: Task
>            Reporter: Tim Allison
>            Priority: Major
>         Attachments: 1008690.pdf
>
>
> PDF/UA includes structural markup. We hacked out a standalone handler for 
> this back in 1.x but haven't touched it in years.
>  
> We should modernize our handling of structural tags and eventually consider 
> turning that on by default. That decision will be based on evaluation on 
> 1000s of PDFs. This is not a default switch to be taken lightly.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to