[jira] [Commented] (TIKA-4444) PDFParser shows wrong data in xmp "dc:subject" tag

Peter Hoogendijk (Jira) Tue, 01 Jul 2025 23:58:36 -0700


    [ 
https://issues.apache.org/jira/browse/TIKA-4444?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17987344#comment-17987344
 ]


Peter Hoogendijk commented on TIKA-4444:
----------------------------------------

Ευχαριστώ for working on this. Nowadays too many tools only present an 
interpretation of the underlying data (even Adobe Acrobat). But for some 
purposes the actual bits are needed, if only to be able to determine why 
something doesn't work as expected. But maybe I'm just getting old, I started 
programming 6502, 8080 and Z80 processors with a hex editor and look where we 
are now ;).

> PDFParser shows wrong data in xmp "dc:subject" tag
> --------------------------------------------------
>
>                 Key: TIKA-4444
>                 URL: https://issues.apache.org/jira/browse/TIKA-4444
>             Project: Tika
>          Issue Type: Bug
>          Components: parser
>    Affects Versions: 3.2.0
>         Environment: * Docker container based on python:3-slim
>  * Debian 12.11
>  * Python 3.13.5
>  * openjdk 17.0.15 2025-04-15
>  * tika-server-standard-3.2.0.jar
>  * tika-server-standard-3.2.2-20250624.143628-8.jar
>  * pdfbox-app-3.0.5.jar
>  * PyPDF 5.6.1
>            Reporter: Peter Hoogendijk
>            Assignee: Tilman Hausherr
>            Priority: Major
>              Labels: xmp
>             Fix For: 4.0.0, 3.2.1
>
>         Attachments: lorem-ipsum.pdf, lorem-ipsum.xml
>
>
> The xmp metadata "dc:subject" tag contains the wrong data: it shows a list 
> with the data from the following tags:
>  * pdf:docinfo:subject (from the pdf metadata)
>  * pdf:docinfo:keywords (from the pdf metadata)
>  * pdf:keywords (from the xmp metadata)
> And it is missing the data from the following tags:
>  * dc:subject (from the xmp metadata)
> When looking at the XML for my testfile (see attachments) the xmp metadata 
> contains the correct "dc:subject" and "pdf:keywords" but:
>  * Tika shows the wrong data in "dc:subject" (from the xmp metadata)
>  * Tika does not show "pdf:keywords" (from the xmp metadata)
>  * Tika does not show the actual "dc:subject" (from the xmp metadata)
> This has been tested with tika-server-standard-3.2.0.jar and 
> tika-server-standard-3.2.2-20250624.143628-8.jar



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

[jira] [Commented] (TIKA-4444) PDFParser shows wrong data in xmp "dc:subject" tag

Reply via email to