Jason Harrop created FOP-3340:
---------------------------------
Summary: A CJK ideograph is written to the PDF's ToUnicode as the
Kangxi radical which shares its glyph
Key: FOP-3340
URL: https://issues.apache.org/jira/browse/FOP-3340
Project: FOP
Issue Type: Bug
Components: font/opentype
Affects Versions: 2.11
Reporter: Jason Harrop
With complex script features enabled, CJK text is painted correctly but cannot
be extracted, searched or read out by a screen reader: every ideograph whose
glyph is shared with a Kangxi radical (U+2E80..U+2FDF) comes out as that
radical.
One FO, one font (Source Han Sans CN), stock FOP, only the complex-script flag
changed:
{noformat}
features on 0003=U+2F63(⽣) 0004=U+2F45(⽅) 0005=U+2F08(⼈) 0006=U+FF0C(,)
0007=U+724B(牋)
features off 0003=U+751F(生) 0004=U+65B9(方) 0005=U+4EBA(人) 0006=U+FF0C(,)
0007=U+724B(牋)
{noformat}
The glyphs drawn are identical either way — glyph 3, 4, 5, 6, 7 at x = 0, 12,
24, 36, 48, advance 1 each; only the ToUnicode differs. U+724B, which no
radical shares, is right either way, which is the control.
h4. Cause
{{MultiByteFont.performSubstitution}} runs the font's layout tables over
characters: characters → glyphs → GSUB → {{mapGlyphsToChars}}.
{{mapGlyphsToChars}} takes each glyph's character from
{{findCharacterFromGlyphIndex}}, whose contract is _"if more than one
correspondence exists, then the first one is returned (ordered by
bfentries\[])"_. A CJK font maps a radical and the ideograph it is the radical
of to one glyph — read out of SourceHanSansCN-Medium's cmap, U+2F63 and U+751F
are both glyph 18742, U+2F45 and U+65B9 are both 14819, U+2F08 and U+4EBA are
both 8966. The radical is the lower code point, so the reverse lookup hands the
layout the radical, and that is what reaches the PDF's ToUnicode.
Nothing about the substitution itself is at fault: the round trip through
characters loses the identity of the character even where GSUB left the glyph
untouched.
h4. Fix
{{mapGlyphsToChars}} should prefer the character the glyph came from — the
{{GlyphSequence}} already carries it in the glyph's {{CharAssociation}} —
wherever the substitution left the glyph alone, i.e. the association covers
exactly one character and the font's character map maps that character to this
same glyph:
{noformat}
int gi = gs.getGlyph(i);
int cc = findUnsubstitutedCharacter(gs, i, ca, nc, gi);
if (cc == 0) {
cc = findCharacterFromGlyphIndex(gi);
}
{noformat}
Anything the substitution did produce — a ligature, or a glyph with no
character of its own — still goes through {{findCharacterFromGlyphIndex}}
exactly as before, and a supplementary plane character still comes back as its
surrogate pair.
h4. Test
{{fop-core/src/test/java/org/apache/fop/fonts/MultiByteFontTestCase.java}},
four cases against a hand-built character map holding the three
radical/ideograph pairs above, so no font need be installed:
* an ideograph sharing its glyph with a radical comes back as the ideograph
(fails without the fix: {{expected:<\[生方人]牋> but was:<\[⽣⽅⼈]牋>}}),
* a radical that really was written stays a radical,
* a glyph the substitution did produce is still mapped through the character
map,
* a supplementary plane character still comes back as its surrogate pair.
fop-core's full suite passes with the fix: 3648 tests, 0 failures, 4 skipped.
h4. Patch
In the PR: one commit, two files,
{{fop-core/src/main/java/org/apache/fop/fonts/MultiByteFont.java}} (+40 −1) and
the new test (+140).
--
This message was sent by Atlassian Jira
(v8.20.10#820010)