[jira] [Commented] (TIKA-4442) PDFParser does not list all metadata extracted by PDFBox

Tilman Hausherr (Jira) Tue, 24 Jun 2025 05:12:10 -0700


    [ 
https://issues.apache.org/jira/browse/TIKA-4442?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17985848#comment-17985848
 ]


Tilman Hausherr commented on TIKA-4442:
---------------------------------------

Thank you, my code gets them all. However it's not perfect, I noticed that in 
the files I mentioned above there are several elements for "type" but my code 
only finds one. My mass search code used xmpbox, while the code for tika uses 
jempbox. I'm not even sure if the files are correct, in one of the descriptions 
it was mentioned that several types should be in the same field. I need to run 
another test first which requires a full build.

> PDFParser does not list all metadata extracted by PDFBox
> --------------------------------------------------------
>
>                 Key: TIKA-4442
>                 URL: https://issues.apache.org/jira/browse/TIKA-4442
>             Project: Tika
>          Issue Type: Improvement
>          Components: parser
>    Affects Versions: 3.2.0
>         Environment: * Docker container based on python:3-slim
>  * Debian 12.11
>  * Python 3.13.5
>  * openjdk 17.0.15 2025-04-15
>  * tika-server-standard-3.2.0.jar
>  * pdfbox-app-3.0.5.jar
>  * PyPDF 5.6.1
>            Reporter: Peter Hoogendijk
>            Priority: Major
>         Attachments: lorem-ipsum.pdf
>
>
> While using Apache Tika to extract metadata from PDF files, I found the 
> following XMP metadata entries to be missing:
>  * dc:identifier
>  * dc:language
>  * dc:publisher
>  * dc:relation
>  * dc:source
>  * dc:type
> Python (PyPDF2) and PDFBox (as used by Tika's PDFParser) do show these XMP 
> metadata entries, so I expected Apache Tika to also extract these entries.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

[jira] [Commented] (TIKA-4442) PDFParser does not list all metadata extracted by PDFBox

Reply via email to