Jason Harrop created FOP-3345:
---------------------------------
Summary: ToUnicode publishes a private-use code point for any
glyph substitution produced, so ligatures and Arabic contextual forms are not
searchable, copyable or readable aloud
Key: FOP-3345
URL: https://issues.apache.org/jira/browse/FOP-3345
Project: FOP
Issue Type: Bug
Components: renderer/pdf, font/opentype
Affects Versions: 2.11
Reporter: Jason Harrop
MultiByteFont.mapGlyphsToChars maps the substituted glyph sequence back to
characters one character per glyph. For a glyph substitution left alone it
finds the originating character. For a glyph substitution produced, whose glyph
index no character maps to, it falls through to createPrivateUseMapping, which
mints a code point from U+E000 upward. The painter later hands that code point
to CIDSubset.mapCodePoint, and CIDSubset.getChars is what PDFToUnicodeCMap is
built from. So the private-use code point minted to give the glyph an identity
becomes its published meaning.
That is every ligature (fi, fl, ffi, ft, ti and so on wherever the font has no
cmap entry for the ligature glyph), and every single substitution whose output
glyph has no cmap entry of its own: Arabic initial, medial and final forms in a
font without presentation-form cmap entries, small capitals, old-style figures,
contextual alternates, and each glyph of a ccmp decomposition.
Measured with the 2.11 command line, read back with pdftotext:
{noformat}
Carlito, script="latn", "office affluent fifty flow ti fi"
-> "office affluent fiy flow fi"
(ft and ti minted; fi, fl, ffi, ffl publish U+FB01..FB04 because
Carlito's
cmap happens to map those forms)
Noto Sans Arabic, script="arab", "السلام عليكم"
-> the medial yeh extracts as U+E001 U+E000 (ccmp split it into
uni066E.medi.wide
and twodotshorizontalbelowar), the other letters as presentation forms
{noformat}
A docx4j corpus render of an Arabic probe through 2.11 counted 105 private-use
characters against 308 base letters and 441 presentation forms.
The page looks right, since the glyph and its advance are correct; the defect
shows only when someone selects, searches or reads the text.
h3. Fix
The information is already there: the GlyphSequence carries a CharAssociation
per glyph naming the characters that produced it, and
findUnsubstitutedCharacter reads it and discards it for exactly the substituted
case.
# MultiByteFont.mapGlyphsToChars records, per glyph index, the characters of
the association whenever the glyph was substituted (findUnsubstitutedCharacter
returned nothing). A disjoint association (a ligature with ignored marks
between its components) contributes its sub-intervals only. The character
returned for layout does not change: it is still the glyph's identity, which
findGlyphIndex maps back to the glyph.
# CIDSet gains getUnicodeSequences(): one String per selector, the recorded
characters where there are any and the code point otherwise. CIDSubset and
CIDFull implement it.
# PDFFactory builds the ToUnicode CMap from that instead of getChars().
# PDFToUnicodeCMap takes one destination String per selector. Only a
destination of one code point may join a bfrange; a multi-character destination
is written as a bfchar with a string, which the spec allows.
Two cases a ToUnicode entry cannot express, and how the fix leaves them:
* Several glyphs for one character (a ccmp decomposition replicates the
association onto each output glyph). The first glyph publishes the character;
the second and later keep the private-use code point. An empty destination was
tried and rejected: poppler reads it correctly, but mupdf prints U+FFFD, pdf.js
a space, and PDFium the raw selector as a control character. Publishing the
character twice makes poppler and mupdf insert a space. ActualText per cluster
is the mechanism for this case and is a content-stream change outside this
issue.
* One glyph seen with two different meanings (a shared dotless base in a
decomposing Arabic font). It publishes neither, keeping the private-use code
point, rather than the wrong letter.
After the fix, with the drawn content byte for byte unchanged:
{noformat}
Carlito -> "office affluent fifty flow ti fi" (pdftotext, mupdf,
pdf.js, PDFium)
Noto Sans Arabic -> "السلام عليكم" (the dots glyph is the one
left)
{noformat}
Tests: PDFToUnicodeCMapTestCase and MultiByteFontTestCase extended; the
writer's existing output is pinned by ToUnicodeCharacterisationTestCase before
the change.
Related: the same rewrite of PDFToUnicodeCMap fixes the selector drift after a
surrogate pair, reported separately (a companion issue filed alongside this
one), because the positional char\[] that caused it is what had to go.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)