[
https://issues.apache.org/jira/browse/TIKA-3571?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17533099#comment-17533099
]
Hudson commented on TIKA-3571:
------------------------------
SUCCESS: Integrated in Jenkins build Tika ยป tika-main-jdk8 #545 (See
[https://ci-builds.apache.org/job/Tika/job/tika-main-jdk8/545/])
TIKA-3571 -- rollback puppycrawl -- requires java > 8 (tallison:
[https://github.com/apache/tika/commit/b4c1c033f2890b5e263ad395477336c9e4e90ee3])
* (edit) tika-parent/pom.xml
> Add an interface for rendering engines
> --------------------------------------
>
> Key: TIKA-3571
> URL: https://issues.apache.org/jira/browse/TIKA-3571
> Project: Tika
> Issue Type: Wish
> Reporter: Tim Allison
> Priority: Major
> Fix For: 2.4.1
>
>
> We've now seen a few requests for extracting text _and_ rendering PDFs, and
> certainly it might be useful to have alternatives for rendering files (e.g.
> this [Alfresco
> study|https://hub.alfresco.com/t5/alfresco-content-services-blog/pdf-rendering-engine-performance-and-fidelity-comparison/ba-p/287618]),
> including MSOffice or at least PPTx...
> And there are cases where users don't want the rendered images, but they do
> want OCR to be run against the rendered images.
> I doubt I'll have a chance to work on this for a while, but I wanted to open
> an issue for discussion.
--
This message was sent by Atlassian Jira
(v8.20.7#820007)