[ 
https://issues.apache.org/jira/browse/TIKA-4891?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18115349#comment-18115349
 ] 

Tim Allison commented on TIKA-4891:
-----------------------------------

I merged this for now. We'll see what the regression run for 4.1.0 shows on the 
1.2 million files. We did run this on ~10k files during development.

[~tilman] , I _think_ I fixed the issue you mentioned. I couldn't reproduce the 
encoding problem... is there a chance the encoding was altered when you 
modified the tags?

> Improve handling of PDF/UA structural tags/marked content
> ---------------------------------------------------------
>
>                 Key: TIKA-4891
>                 URL: https://issues.apache.org/jira/browse/TIKA-4891
>             Project: Tika
>          Issue Type: Task
>            Reporter: Tim Allison
>            Priority: Major
>         Attachments: 1008690.pdf, TIKA-4891-1008690.html, 
> image-2026-09-13-10-54-59-251.png, screenshot-1.png, screenshot-2.png
>
>
> PDF/UA includes structural markup. We hacked out a standalone handler for 
> this back in 1.x but haven't touched it in years.
>  
> We should modernize our handling of structural tags and eventually consider 
> turning that on by default. That decision will be based on evaluation on 
> 1000s of PDFs. This is not a default switch to be taken lightly.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to