[
https://issues.apache.org/jira/browse/FOP-3347?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18124628#comment-18124628
]
Jason Harrop commented on FOP-3347:
-----------------------------------
Yes. Attached: FOP-3347-format-character.fo and fop-test-fonts.xconf, which
declares DejaVuLGCSerif.ttf from the test tree; it maps U+206A (and U+200E,
U+200F, U+2060, U+200B) to a glyph of advance 0. Copy the config to the root of
a checkout and run from there:
{noformat}
fop -c fop-test-fonts.xconf -fo FOP-3347-format-character.fo -pdf out.pdf
pdftotext out.pdf -
{noformat}
The FO is one block, {{abcdef ghi}} in DejaVuLGCSerif.
On main at 5be8c69b6 pdftotext gives {{abcdef ghi}} and the ToUnicode CMap has
no entry for U+206A: a single range {{<0003> <000b> <0061>}} covers a to i.
With the pull request's change pdftotext gives {{abcdef ghi}}, the U+206A
between c and d, and the CMap is
{noformat}
<0003> <0005> <0061> <0006> <206a> <0007> <000c> <0064>
{noformat}
Nothing drawn changes: pdftotext -bbox gives the same two word boxes either way
(56.69 to 97.46 and 101.26 to 120.50 pt).
> Directional marks and other format characters are dropped from the PDF text
> layer for any font that goes through the complex-scripts mapping, even when
> the font has a zero-width glyph for them
> ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
>
> Key: FOP-3347
> URL: https://issues.apache.org/jira/browse/FOP-3347
> Project: FOP
> Issue Type: Bug
> Components: font/opentype, renderer/pdf
> Affects Versions: 2.11
> Reporter: Jason Harrop
> Priority: Minor
> Attachments: FOP-3347-format-character.fo, fop-test-fonts.xconf
>
>
> MultiByteFont.performSubstitution maps characters to glyphs, runs GSUB, and
> then, unless retainControls (which TextLayoutManager never sets), calls
> elideControls, which removes every glyph whose characters are all "elidable
> controls": C0 and C1 controls, U+200B to U+200F, U+2028 to U+202E, U+2060,
> U+2066 to U+206F. The glyph and its association are gone, so the character
> never reaches the CID subset or the ToUnicode CMap. Text extraction, search
> and screen readers lose it.
> The elision is right for a font with no glyph for the character, which would
> otherwise be drawn as the missing-character glyph. But common fonts carry a
> real, zero-width glyph for exactly these characters: Arimo, Tinos and DejaVu
> Sans map U+200E, U+200F and U+206A to glyphs of advance 0; Carlito maps
> U+200E and U+200F. For those the character can simply stay. The single-byte
> path keeps it already (it maps the character and draws its zero-width glyph),
> so the same document extracts differently depending on which encoding mode
> the font was declared with. Word's PDF export keeps the marks.
> Reproducer, 2.11 command line, Arimo with the default (CID) declaration:
> {noformat}
> <fo:block font-family="Arimo">abcdef ghi</fo:block>
> {noformat}
> with U+206A in place of the mark, pdftotext gives "abcdef" and the ToUnicode
> CMap has no entry for it. On a corpus of 598 documents converted from Word,
> 19 U+206A disappeared this way where Word's own PDFs keep them (and 67 bidi
> marks, which Word drops too; see the fix).
> h3. Fix
> In elideControls, keep an association of exactly one format character (U+2000
> to U+206F) whose glyph is the one the character map gives it and whose
> advance in this font is zero, unless it is a bidi control (U+200E, U+200F,
> U+202A-202E, U+2066-2069). It costs no space, draws nothing, and gets its own
> subset selector and ToUnicode entry. A character the font has no glyph for, a
> glyph with an advance (a font defect: Tinos maps U+2060 to a 799-unit glyph),
> a bidi control and the C0 and C1 controls are elided as before.
> The bidi controls stay out on purpose. The text layer is written in visual
> order, after FOP's own bidi resolution, so a control there has done its work
> and a reader applies it a second time: a first cut that kept them had
> pdftotext wrap every kept mark in U+202B and U+202C. Word's PDF export drops
> them too (measured on 597 documents: its text layer holds the U+206A of two
> documents and none of the 66 right-to-left marks of another).
> After the fix, "abc<U+206A>def" extracts with the U+206A and every word box
> is identical (56.69-92.70, 96.03-112.04 pt at 12pt); the U+200F reproducer is
> unchanged, by design.
> One consequence to know: a kept format glyph is in the sequence during GSUB
> and GPOS, so a ZWNJ between f and i now blocks the ligature (its purpose),
> and a format character between two kerned letters blocks the pair, where a
> shaper that treats default-ignorables as invisible would not. Marks at word
> edges, the common case, touch nothing.
> Test: MultiByteFontTestCase, five cases on a hand-built character map and
> width array: a zero-width U+206A kept, a zero-width U+200F elided as a bidi
> control, one with no glyph elided, one with an advance elided, a C0 control
> elided.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)