[
https://issues.apache.org/jira/browse/TIKA-4883?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Tim Allison resolved TIKA-4883.
-------------------------------
Fix Version/s: 4.1.0
Resolution: Fixed
> Don't dump raw text when OCR is triggered in AUTO mode in PDFs
> --------------------------------------------------------------
>
> Key: TIKA-4883
> URL: https://issues.apache.org/jira/browse/TIKA-4883
> Project: Tika
> Issue Type: Task
> Reporter: Tim Allison
> Priority: Minor
> Fix For: 4.1.0
>
>
> In AUTO OCR mode in PDFs, if there's not enough content or if there are a lot
> of characters missing unicode mappings, we trigger OCR on that page.
> In the current implementation, we write that content to the page, and then
> make the determination. If OCR is triggered, we append the OCR content to the
> end of the page.
> It would be cleaner in AUTO mode, if we cached the page content, ran the
> score and then decided whether or not to run OCR WITHOUT including the
> original page content before making that determination.
> Or, in short, don't dump trash then concatenate OCR... go only with OCR in
> AUTO mode, when triggered.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)