[ 
https://issues.apache.org/jira/browse/TIKA-4946?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18121960#comment-18121960
 ] 

Tim Allison edited comment on TIKA-4946 at 10/2/26 12:32 PM:
-------------------------------------------------------------

Claude recommends the following. If we have anyone with actual image processing 
experience who wants to chime in, please do.

 
{noformat}
1. Contrast check (standard deviation)
A blank page is nearly uniform, so its grayscale standard deviation is tiny.
bashmagick page.png -colorspace Gray -format "%[fx:standard_deviation]\n" info:
Blank scans usually come out well under ~0.02, and text pages are much higher. 
Shave the borders first (-shave 3%x3%) so scanner edges, shadows, and punch 
holes don't inflate the number.
2. Ink ratio (fraction of dark pixels)
bashmagick page.png -colorspace Gray -shave 3%x3% -median 3 \
  -threshold 60% -negate -format "%[fx:mean]\n" info:
-median 3 (or -despeck) removes speckle noise before thresholding. Use a fixed 
threshold here, not -auto-threshold otsu. On a uniform page, Otsu still splits 
the noise into two classes, and a blank sheet can come out as ~50% "ink." This 
one bites people.
3. Trim test
Trim away everything close to the background color and see what's left.
bashmagick page.png -shave 3%x3% -fuzz 15% -trim -format "%w %h\n" info:
A blank page collapses to something tiny, like 1 1, while a page with content 
keeps a real bounding box. This is quick and surprisingly effective.
4. Connected components (text-likeness)
Text produces many small blobs; noise produces a few specks.
bashmagick page.png -colorspace Gray -shave 3%x3% -threshold 60% -negate \
  -define connected-components:verbose=true \
  -define connected-components:area-threshold=15 \
  -connected-components 8 null: | tail -n +2 | wc -l
A real text page typically has hundreds or thousands of components, and a blank 
one has a handful. This is the best ImageMagick-only discriminator, because it 
also catches pages that have marks but no text (a stray line, a stamp edge). 
{noformat}


was (Author: [email protected]):
Claude recommends the following. If we have anyone with actual image processing 
experience who wants to chime in, please do.

 
{noformat}
1. Contrast check (standard deviation)
A blank page is nearly uniform, so its grayscale standard deviation is tiny.
bashmagick page.png -colorspace Gray -format "%[fx:standard_deviation]\n" info:
Blank scans usually come out well under ~0.02, and text pages are much higher. 
Shave the borders first (-shave 3%x3%) so scanner edges, shadows, and punch 
holes don't inflate the number.
2. Ink ratio (fraction of dark pixels)
bashmagick page.png -colorspace Gray -shave 3%x3% -median 3 \
  -threshold 60% -negate -format "%[fx:mean]\n" info:
-median 3 (or -despeck) removes speckle noise before thresholding. Use a fixed 
threshold here, not -auto-threshold otsu. On a uniform page, Otsu still splits 
the noise into two classes, and a blank sheet can come out as ~50% "ink." This 
one bites people.
3. Trim test
Trim away everything close to the background color and see what's left.
bashmagick page.png -shave 3%x3% -fuzz 15% -trim -format "%w %h\n" info:
A blank page collapses to something tiny, like 1 1, while a page with content 
keeps a real bounding box. This is quick and surprisingly effective.
4. Connected components (text-likeness)
Text produces many small blobs; noise produces a few specks.
bashmagick page.png -colorspace Gray -shave 3%x3% -threshold 60% -negate \
  -define connected-components:verbose=true \
  -define connected-components:area-threshold=15 \
  -connected-components 8 null: | tail -n +2 | wc -l
A real text page typically has hundreds or thousands of components, and a blank 
one has a handful. This is the best ImageMagick-only discriminator, because it 
also catches pages that have marks but no text (a stray line, a stamp edge). 
{noformat}

> Don't send blank pages for OCR
> ------------------------------
>
>                 Key: TIKA-4946
>                 URL: https://issues.apache.org/jira/browse/TIKA-4946
>             Project: Tika
>          Issue Type: Task
>            Reporter: Tim Allison
>            Priority: Major
>
> Some vlm-based OCR tools will hallucinate an entire journal article page's 
> worth of content if we send a blank image.
> Ideally, the ocr tools would prevent this, but we should try to do something 
> on our side to prevent this.
> Not sure where to put this: in the PDFParser or in the OCR engines?



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to