[ 
https://issues.apache.org/jira/browse/TIKA-4856?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18109474#comment-18109474
 ] 

ASF GitHub Bot commented on TIKA-4856:
--------------------------------------

dschmidt opened a new pull request, #3096:
URL: https://github.com/apache/tika/pull/3096

   Getting "the thumbnail of this file" out of Tika takes format knowledge on 
the client: the THUMBNAIL of a Word or Excel file is an EMF/WMF whose usable 
form is the RENDERING underneath it, a PDF has no THUMBNAIL but a page 
RENDERING, the THUMBNAIL of a DOCX inside a ZIP is not the ZIP's, and with 
rendering enabled the picture of an embedded OLE object is a RENDERING too, 
plus the request configuration that switches the renderers on.
   
   `POST`/`PUT /unpack/thumbnail`, next to `/unpack` and `/unpack/all`, answers 
with the thumbnail as JSON: the `/rmeta` metadata object of the embedded 
document that is the thumbnail and the image as base64, or `204` if the 
document has none.
   
   ```
   {
     "metadata": { "Content-Type": "image/png", "tiff:ImageWidth": "800", 
"tk:embedded-resource-type": "RENDERING", ... },
     "image": "iVBORw0KGgo..."
   }
   ```
   
   It is one forked parse in unpack mode with a fixed parse context: text 
extraction off, PDF page 1 rendered at 96 dpi without OCR, EMF/WMF thumbnails 
rendered (`renderOnlyEmbeddedResourceTypes: ["THUMBNAIL"]`, from #3095, so the 
pictures of embedded objects are not rendered), OCR of images off, only 
THUMBNAIL and RENDERING embedded documents extracted, together with their 
metadata. `ThumbnailSelector` then picks, in this order, the raster THUMBNAIL 
directly below the document, the rendering of a vector THUMBNAIL, or the 
RENDERING of the first page. It extracts what the document carries; it does not 
resize or convert. A request-supplied parser configuration wins over the fixed 
one.
   
   Verified against a server built from main plus the open thumbnail PRs: docx, 
xlsx, doc, xls, ppt, pptx, odt, epub, GeoGebra, Pages, Numbers, Keynote, mp3, 
m4a, flac, ogg, pdf, nef and pef all answer with the right image; a zip, a 
plain jpeg and a doc without a thumbnail answer 204. Raw camera files need 
their file name (`Content-Disposition`), as they are detected by extension.
   
   Note: `embedded-limits.maxDepth` is 3 in the fixed context because of the 
off-by-one in the depth limit (separate ticket); 2 would be correct once that 
is fixed.
   
   https://issues.apache.org/jira/browse/TIKA-4856
   




> /unpack/thumbnail: return the document thumbnail with its metadata
> ------------------------------------------------------------------
>
>                 Key: TIKA-4856
>                 URL: https://issues.apache.org/jira/browse/TIKA-4856
>             Project: Tika
>          Issue Type: New Feature
>            Reporter: Dominik Schmidt
>            Priority: Major
>
> With TIKA-4850 through TIKA-4855 every container format that carries a 
> thumbnail emits it as a THUMBNAIL embedded document, the PDF parser renders 
> pages as RENDERING documents, and the EMF/WMF renderer turns the vector 
> thumbnails of Office documents into raster ones. Getting "the thumbnail of 
> this file" out of that still takes format knowledge on the client: the 
> THUMBNAIL of a Word or Excel file is an EMF/WMF whose usable form is the 
> RENDERING underneath it, a PDF has no THUMBNAIL but a page RENDERING, the 
> THUMBNAIL of a DOCX inside a ZIP is not the ZIP's, and with rendering enabled 
> the picture of an embedded OLE object is a RENDERING too. Plus the request 
> config that switches the renderers on.
> Proposal: POST /unpack/thumbnail next to /unpack and /unpack/all, multipart 
> like them. It runs the usual forked parse in unpack mode with a fixed parse 
> context (PDF page 1 rendered, EMF/WMF rendered) and picks, in this order: the 
> raster THUMBNAIL at depth 1; the rendering of that thumbnail; the depth-1 
> RENDERING of PDF page 1. The endpoint extracts what the document carries; it 
> does not resize, convert or generate previews.
> The response is JSON: the /rmeta metadata object of the selected embedded 
> document, and the image as base64. Thumbnails are small, so the encoding 
> overhead does not matter, and the caller gets type, dimensions, origin 
> (stored thumbnail or rendering, tk:rendering:rendered-by) and path in one 
> round trip without unpacking a zip. 204 when the document has no thumbnail.
> {
>   "metadata": {
>     "Content-Type": "image/png",
>     "Content-Length": "8459",
>     "tiff:ImageWidth": "800",
>     "tiff:ImageLength": "1131",
>     "tk:embedded-resource-type": "RENDERING",
>     "tk:embedded-resource-path": "/thumbnail.emf/thumbnail.png",
>     "tk:embedded-depth": "2",
>     "tk:rendering:rendered-by": "poi-metafile-renderer",
>     "tk:resource-name": "thumbnail.png"
>   },
>   "image": "iVBORw0KGgoAAAANSUhEUgAA..."
> }
> To keep the selection rule short, the metafile renderer could give the 
> rendering of a THUMBNAIL the THUMBNAIL type as well (its 
> tk:rendering:rendered-by tells it apart), so a raster thumbnail is a 
> THUMBNAIL regardless of whether the document stored it as PNG or as EMF.
> What do you think? 



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to