[
https://issues.apache.org/jira/browse/PDFBOX-6256?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18119826#comment-18119826
]
Tilman Hausherr commented on PDFBOX-6256:
-----------------------------------------
Re inline, see the file from PDFBOX-3248 and PDFBOX-5868 (or the files created
after the change). I hadn't bothered to look at the content stream until now
because the tests looked good and passed. But copilot is right. This is a
similar problem to what we had with MCID in PDFBOX-5890. We could create
another beginMarkedContent() method with a string, but I'm not really sure, and
we can still do it later.
> Extracted Text is incorrect. Option for ActualText proposed.
> ------------------------------------------------------------
>
> Key: PDFBOX-6256
> URL: https://issues.apache.org/jira/browse/PDFBOX-6256
> Project: PDFBox
> Issue Type: Bug
> Affects Versions: 3.0.8 PDFBox
> Reporter: Volker Kunert
> Priority: Major
> Fix For: 3.0.9 PDFBox, 4.0.0
>
>
> For a correct result of text extraction an option to use ActualText is added
> to GlyphLayoutProcessor(Awt|Fop)
> SeeĀ https://github.com/apache/pdfbox/pull/526
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]