Jason Harrop created FOP-3340:
---------------------------------

             Summary: A CJK ideograph is written to the PDF's ToUnicode as the 
Kangxi radical which shares its glyph
                 Key: FOP-3340
                 URL: https://issues.apache.org/jira/browse/FOP-3340
             Project: FOP
          Issue Type: Bug
          Components: font/opentype
    Affects Versions: 2.11
            Reporter: Jason Harrop


With complex script features enabled, CJK text is painted correctly but cannot 
be extracted, searched or read out by a screen reader: every ideograph whose 
glyph is shared with a Kangxi radical (U+2E80..U+2FDF) comes out as that 
radical.

One FO, one font (Source Han Sans CN), stock FOP, only the complex-script flag 
changed:

{noformat}
features on   0003=U+2F63(⽣) 0004=U+2F45(⽅) 0005=U+2F08(⼈) 0006=U+FF0C(,) 
0007=U+724B(牋)
features off  0003=U+751F(生) 0004=U+65B9(方) 0005=U+4EBA(人) 0006=U+FF0C(,) 
0007=U+724B(牋)
{noformat}

The glyphs drawn are identical either way — glyph 3, 4, 5, 6, 7 at x = 0, 12, 
24, 36, 48, advance 1 each; only the ToUnicode differs. U+724B, which no 
radical shares, is right either way, which is the control.

h4. Cause

{{MultiByteFont.performSubstitution}} runs the font's layout tables over 
characters: characters → glyphs → GSUB → {{mapGlyphsToChars}}. 
{{mapGlyphsToChars}} takes each glyph's character from 
{{findCharacterFromGlyphIndex}}, whose contract is _"if more than one 
correspondence exists, then the first one is returned (ordered by 
bfentries\[])"_. A CJK font maps a radical and the ideograph it is the radical 
of to one glyph — read out of SourceHanSansCN-Medium's cmap, U+2F63 and U+751F 
are both glyph 18742, U+2F45 and U+65B9 are both 14819, U+2F08 and U+4EBA are 
both 8966. The radical is the lower code point, so the reverse lookup hands the 
layout the radical, and that is what reaches the PDF's ToUnicode.

Nothing about the substitution itself is at fault: the round trip through 
characters loses the identity of the character even where GSUB left the glyph 
untouched.

h4. Fix

{{mapGlyphsToChars}} should prefer the character the glyph came from — the 
{{GlyphSequence}} already carries it in the glyph's {{CharAssociation}} — 
wherever the substitution left the glyph alone, i.e. the association covers 
exactly one character and the font's character map maps that character to this 
same glyph:

{noformat}
int gi = gs.getGlyph(i);
int cc = findUnsubstitutedCharacter(gs, i, ca, nc, gi);
if (cc == 0) {
    cc = findCharacterFromGlyphIndex(gi);
}
{noformat}

Anything the substitution did produce — a ligature, or a glyph with no 
character of its own — still goes through {{findCharacterFromGlyphIndex}} 
exactly as before, and a supplementary plane character still comes back as its 
surrogate pair.

h4. Test

{{fop-core/src/test/java/org/apache/fop/fonts/MultiByteFontTestCase.java}}, 
four cases against a hand-built character map holding the three 
radical/ideograph pairs above, so no font need be installed:

* an ideograph sharing its glyph with a radical comes back as the ideograph 
(fails without the fix: {{expected:<\[生方人]牋> but was:<\[⽣⽅⼈]牋>}}),
* a radical that really was written stays a radical,
* a glyph the substitution did produce is still mapped through the character 
map,
* a supplementary plane character still comes back as its surrogate pair.

fop-core's full suite passes with the fix: 3648 tests, 0 failures, 4 skipped.

h4. Patch

In the PR: one commit, two files, 
{{fop-core/src/main/java/org/apache/fop/fonts/MultiByteFont.java}} (+40 −1) and 
the new test (+140).




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to