[
https://issues.apache.org/jira/browse/FOP-3340?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18124629#comment-18124629
]
Jason Harrop commented on FOP-3340:
-----------------------------------
Yes, two FOs. The mechanism is not CJK-specific: it is any two code points the
font's cmap maps to one glyph, and the test tree has such a font, so the first
needs nothing installed.
Aegean600.ttf maps U+2019 (right single quotation mark) and U+02BC (modifier
letter apostrophe) to one glyph, and U+2018 and U+02BB likewise. Attached:
FOP-3340-shared-glyph.fo, one block, {{it’s}} in Aegean600, and
fop-test-fonts.xconf. Copy the config to the root of a checkout and run from
there:
{noformat}
fop -c fop-test-fonts.xconf -fo FOP-3340-shared-glyph.fo -pdf out.pdf
pdftotext out.pdf -
{noformat}
On main at 5be8c69b6 pdftotext gives {{itʼs}} with U+02BC, and the ToUnicode
CMap has {{<0005> <02bc>}}. With the pull request's change it gives {{it’s}}
and {{<0005> <2019>}}.
The CJK case of the description is FOP-3340-cjk-radical.fo, also attached: one
block, {{生方人,牋}} in font-family "Source Han Sans CN" (Adobe, OFL,
https://github.com/adobe-fonts/source-han-sans). Add to the config
{noformat}
<font embed-url="file:///path/to/SourceHanSansCN-Regular.otf"><font-triplet
name="Source Han Sans CN" style="normal" weight="normal"/></font>
{noformat}
On main pdftotext gives {{⽣⽅⼈,牋}}, the Kangxi radicals U+2F63 U+2F45 U+2F08;
with the change {{生方人,牋}}, U+751F U+65B9 U+4EBA. U+FF0C and U+724B are right
either way, the control.
On the other tickets: FOP-3354 and FOP-3355 carry layout test cases in their
pull requests, as FOP-3348 did; FOP-2918 a layout test case and an FO under
fop/test/xml/pdf-encoding. I will add an FO to each of the remaining font
tickets in turn.
> A CJK ideograph is written to the PDF's ToUnicode as the Kangxi radical which
> shares its glyph
> ----------------------------------------------------------------------------------------------
>
> Key: FOP-3340
> URL: https://issues.apache.org/jira/browse/FOP-3340
> Project: FOP
> Issue Type: Bug
> Components: font/opentype
> Affects Versions: 2.11
> Reporter: Jason Harrop
> Priority: Minor
> Attachments: FOP-3340-cjk-radical.fo, FOP-3340-shared-glyph.fo,
> fop-test-fonts.xconf
>
>
> With complex script features enabled, CJK text is painted correctly but
> cannot be extracted, searched or read out by a screen reader: every ideograph
> whose glyph is shared with a Kangxi radical (U+2E80..U+2FDF) comes out as
> that radical.
> One FO, one font (Source Han Sans CN), stock FOP, only the complex-script
> flag changed:
> {noformat}
> features on 0003=U+2F63(⽣) 0004=U+2F45(⽅) 0005=U+2F08(⼈) 0006=U+FF0C(,)
> 0007=U+724B(牋)
> features off 0003=U+751F(生) 0004=U+65B9(方) 0005=U+4EBA(人) 0006=U+FF0C(,)
> 0007=U+724B(牋)
> {noformat}
> The glyphs drawn are identical either way — glyph 3, 4, 5, 6, 7 at x = 0, 12,
> 24, 36, 48, advance 1 each; only the ToUnicode differs. U+724B, which no
> radical shares, is right either way, which is the control.
> h4. Cause
> {{MultiByteFont.performSubstitution}} runs the font's layout tables over
> characters: characters → glyphs → GSUB → {{mapGlyphsToChars}}.
> {{mapGlyphsToChars}} takes each glyph's character from
> {{findCharacterFromGlyphIndex}}, whose contract is _"if more than one
> correspondence exists, then the first one is returned (ordered by
> bfentries\[])"_. A CJK font maps a radical and the ideograph it is the
> radical of to one glyph — read out of SourceHanSansCN-Medium's cmap, U+2F63
> and U+751F are both glyph 18742, U+2F45 and U+65B9 are both 14819, U+2F08 and
> U+4EBA are both 8966. The radical is the lower code point, so the reverse
> lookup hands the layout the radical, and that is what reaches the PDF's
> ToUnicode.
> Nothing about the substitution itself is at fault: the round trip through
> characters loses the identity of the character even where GSUB left the glyph
> untouched.
> h4. Fix
> {{mapGlyphsToChars}} should prefer the character the glyph came from — the
> {{GlyphSequence}} already carries it in the glyph's {{CharAssociation}} —
> wherever the substitution left the glyph alone, i.e. the association covers
> exactly one character and the font's character map maps that character to
> this same glyph:
> {noformat}
> int gi = gs.getGlyph(i);
> int cc = findUnsubstitutedCharacter(gs, i, ca, nc, gi);
> if (cc == 0) {
> cc = findCharacterFromGlyphIndex(gi);
> }
> {noformat}
> Anything the substitution did produce — a ligature, or a glyph with no
> character of its own — still goes through {{findCharacterFromGlyphIndex}}
> exactly as before, and a supplementary plane character still comes back as
> its surrogate pair.
> h4. Test
> {{fop-core/src/test/java/org/apache/fop/fonts/MultiByteFontTestCase.java}},
> four cases against a hand-built character map holding the three
> radical/ideograph pairs above, so no font need be installed:
> * an ideograph sharing its glyph with a radical comes back as the ideograph
> (fails without the fix: {{expected:<\[生方人]牋> but was:<\[⽣⽅⼈]牋>}}),
> * a radical that really was written stays a radical,
> * a glyph the substitution did produce is still mapped through the character
> map,
> * a supplementary plane character still comes back as its surrogate pair.
> fop-core's full suite passes with the fix: 3648 tests, 0 failures, 4 skipped.
> h4. Patch
> In the PR: one commit, two files,
> {{fop-core/src/main/java/org/apache/fop/fonts/MultiByteFont.java}} (+40 −1)
> and the new test (+140).
--
This message was sent by Atlassian Jira
(v8.20.10#820010)