Hi,

Sorry about that. I've been looking at this again. My thought is that the first priority is to make sure that correct files work, so I've remove some of the workarounds, and in the worst case some of the older (broken) files won't work anymore. You'll still have to wait for a release. Until then, build from source and remove this call:
https://github.com/apache/pdfbox/blob/2.0.28/pdfbox/src/main/java/org/apache/pdfbox/pdfparser/BaseParser.java#L480
the line is likely to be different in the current source, but the point is to remove "braces = checkForEndOfString(braces);" after the "// PDFBox 276" comment.

Tilman

Am 05.10.2026 um 13:17 schrieb Hyoga Kono:
Hi,

I ran into this issue with PDFBox 3.0.8 as well.

The byte sequence 5C 29 0A 3E appears in a literal string in the /ID array of an xref stream.
PDFBox fails to parse the xref stream, and I can't read the Info metadata.

I've attached a small reproducer using synthetic PDFs.
The literal-string version fails, while the hex-string version with identical bytes works. To run the reproducer, place all three files in the same directory along with pdfbox-app-3.0.8.jar, then run:

```
java --class-path pdfbox-app-3.0.8.jar Reproduce.java
```


Has there been any progress on PDFBOX-6167?
Is there a workaround on the PDFBox side?

Thanks,
Hyoga

On 2026/02/17 15:18:10 Tilman Hausherr wrote:
> Thank you, I've created https://issues.apache.org/jira/browse/PDFBOX-6167
>
> Tilman
>

---------------------------------------------------------------------
To unsubscribe, e-mail:[email protected]
For additional commands, e-mail:[email protected]

Reply via email to