Hi,
Sorry about that. I've been looking at this again. My thought is that
the first priority is to make sure that correct files work, so I've
remove some of the workarounds, and in the worst case some of the older
(broken) files won't work anymore. You'll still have to wait for a
release. Until then, build from source and remove this call:
https://github.com/apache/pdfbox/blob/2.0.28/pdfbox/src/main/java/org/apache/pdfbox/pdfparser/BaseParser.java#L480
the line is likely to be different in the current source, but the point
is to remove "braces = checkForEndOfString(braces);" after the "//
PDFBox 276" comment.
Tilman
Am 05.10.2026 um 13:17 schrieb Hyoga Kono:
Hi,
I ran into this issue with PDFBox 3.0.8 as well.
The byte sequence 5C 29 0A 3E appears in a literal string in the /ID
array of an xref stream.
PDFBox fails to parse the xref stream, and I can't read the Info metadata.
I've attached a small reproducer using synthetic PDFs.
The literal-string version fails, while the hex-string version with
identical bytes works.
To run the reproducer, place all three files in the same directory
along with pdfbox-app-3.0.8.jar, then run:
```
java --class-path pdfbox-app-3.0.8.jar Reproduce.java
```
Has there been any progress on PDFBOX-6167?
Is there a workaround on the PDFBox side?
Thanks,
Hyoga
On 2026/02/17 15:18:10 Tilman Hausherr wrote:
> Thank you, I've created
https://issues.apache.org/jira/browse/PDFBOX-6167
>
> Tilman
>
---------------------------------------------------------------------
To unsubscribe, e-mail:[email protected]
For additional commands, e-mail:[email protected]