Jason Harrop created FOP-3347:
---------------------------------
Summary: Directional marks and other format characters are dropped
from the PDF text layer for any font that goes through the complex-scripts
mapping, even when the font has a zero-width glyph for them
Key: FOP-3347
URL: https://issues.apache.org/jira/browse/FOP-3347
Project: FOP
Issue Type: Bug
Components: renderer/pdf, font/opentype
Affects Versions: 2.11
Reporter: Jason Harrop
MultiByteFont.performSubstitution maps characters to glyphs, runs GSUB, and
then, unless retainControls (which TextLayoutManager never sets), calls
elideControls, which removes every glyph whose characters are all "elidable
controls": C0 and C1 controls, U+200B to U+200F, U+2028 to U+202E, U+2060,
U+2066 to U+206F. The glyph and its association are gone, so the character
never reaches the CID subset or the ToUnicode CMap. Text extraction, search and
screen readers lose it.
The elision is right for a font with no glyph for the character, which would
otherwise be drawn as the missing-character glyph. But common fonts carry a
real, zero-width glyph for exactly these characters: Arimo, Tinos and DejaVu
Sans map U+200E, U+200F and U+206A to glyphs of advance 0; Carlito maps U+200E
and U+200F. For those the character can simply stay. The single-byte path keeps
it already (it maps the character and draws its zero-width glyph), so the same
document extracts differently depending on which encoding mode the font was
declared with. Word's PDF export keeps the marks.
Reproducer, 2.11 command line, Arimo with the default (CID) declaration:
{noformat}
<fo:block font-family="Arimo">abcdef ghi</fo:block>
{noformat}
with U+206A in place of the mark, pdftotext gives "abcdef" and the ToUnicode
CMap has no entry for it. On a corpus of 598 documents converted from Word, 19
U+206A disappeared this way where Word's own PDFs keep them (and 67 bidi marks,
which Word drops too; see the fix).
h3. Fix
In elideControls, keep an association of exactly one format character (U+2000
to U+206F) whose glyph is the one the character map gives it and whose advance
in this font is zero, unless it is a bidi control (U+200E, U+200F, U+202A-202E,
U+2066-2069). It costs no space, draws nothing, and gets its own subset
selector and ToUnicode entry. A character the font has no glyph for, a glyph
with an advance (a font defect: Tinos maps U+2060 to a 799-unit glyph), a bidi
control and the C0 and C1 controls are elided as before.
The bidi controls stay out on purpose. The text layer is written in visual
order, after FOP's own bidi resolution, so a control there has done its work
and a reader applies it a second time: a first cut that kept them had pdftotext
wrap every kept mark in U+202B and U+202C. Word's PDF export drops them too
(measured on 597 documents: its text layer holds the U+206A of two documents
and none of the 66 right-to-left marks of another).
After the fix, "abc<U+206A>def" extracts with the U+206A and every word box is
identical (56.69-92.70, 96.03-112.04 pt at 12pt); the U+200F reproducer is
unchanged, by design.
One consequence to know: a kept format glyph is in the sequence during GSUB and
GPOS, so a ZWNJ between f and i now blocks the ligature (its purpose), and a
format character between two kerned letters blocks the pair, where a shaper
that treats default-ignorables as invisible would not. Marks at word edges, the
common case, touch nothing.
Test: MultiByteFontTestCase, five cases on a hand-built character map and width
array: a zero-width U+206A kept, a zero-width U+200F elided as a bidi control,
one with no glyph elided, one with an advance elided, a C0 control elided.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)