[ 
https://issues.apache.org/jira/browse/FOP-3340?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18124629#comment-18124629
 ] 

Jason Harrop commented on FOP-3340:
-----------------------------------

Yes, two FOs. The mechanism is not CJK-specific: it is any two code points the 
font's cmap maps to one glyph, and the test tree has such a font, so the first 
needs nothing installed.
Aegean600.ttf maps U+2019 (right single quotation mark) and U+02BC (modifier 
letter apostrophe) to one glyph, and U+2018 and U+02BB likewise. Attached: 
FOP-3340-shared-glyph.fo, one block, {{it’s}} in Aegean600, and 
fop-test-fonts.xconf. Copy the config to the root of a checkout and run from 
there:
{noformat}
fop -c fop-test-fonts.xconf -fo FOP-3340-shared-glyph.fo -pdf out.pdf
pdftotext out.pdf -
{noformat}
On main at 5be8c69b6 pdftotext gives {{itʼs}} with U+02BC, and the ToUnicode 
CMap has {{<0005> <02bc>}}. With the pull request's change it gives {{it’s}} 
and {{<0005> <2019>}}.
The CJK case of the description is FOP-3340-cjk-radical.fo, also attached: one 
block, {{生方人,牋}} in font-family "Source Han Sans CN" (Adobe, OFL, 
https://github.com/adobe-fonts/source-han-sans). Add to the config
{noformat}
<font embed-url="file:///path/to/SourceHanSansCN-Regular.otf"><font-triplet 
name="Source Han Sans CN" style="normal" weight="normal"/></font>
{noformat}
On main pdftotext gives {{⽣⽅⼈,牋}}, the Kangxi radicals U+2F63 U+2F45 U+2F08; 
with the change {{生方人,牋}}, U+751F U+65B9 U+4EBA. U+FF0C and U+724B are right 
either way, the control.
On the other tickets: FOP-3354 and FOP-3355 carry layout test cases in their 
pull requests, as FOP-3348 did; FOP-2918 a layout test case and an FO under 
fop/test/xml/pdf-encoding. I will add an FO to each of the remaining font 
tickets in turn.

> A CJK ideograph is written to the PDF's ToUnicode as the Kangxi radical which 
> shares its glyph
> ----------------------------------------------------------------------------------------------
>
>                 Key: FOP-3340
>                 URL: https://issues.apache.org/jira/browse/FOP-3340
>             Project: FOP
>          Issue Type: Bug
>          Components: font/opentype
>    Affects Versions: 2.11
>            Reporter: Jason Harrop
>            Priority: Minor
>         Attachments: FOP-3340-cjk-radical.fo, FOP-3340-shared-glyph.fo, 
> fop-test-fonts.xconf
>
>
> With complex script features enabled, CJK text is painted correctly but 
> cannot be extracted, searched or read out by a screen reader: every ideograph 
> whose glyph is shared with a Kangxi radical (U+2E80..U+2FDF) comes out as 
> that radical.
> One FO, one font (Source Han Sans CN), stock FOP, only the complex-script 
> flag changed:
> {noformat}
> features on   0003=U+2F63(⽣) 0004=U+2F45(⽅) 0005=U+2F08(⼈) 0006=U+FF0C(,) 
> 0007=U+724B(牋)
> features off  0003=U+751F(生) 0004=U+65B9(方) 0005=U+4EBA(人) 0006=U+FF0C(,) 
> 0007=U+724B(牋)
> {noformat}
> The glyphs drawn are identical either way — glyph 3, 4, 5, 6, 7 at x = 0, 12, 
> 24, 36, 48, advance 1 each; only the ToUnicode differs. U+724B, which no 
> radical shares, is right either way, which is the control.
> h4. Cause
> {{MultiByteFont.performSubstitution}} runs the font's layout tables over 
> characters: characters → glyphs → GSUB → {{mapGlyphsToChars}}. 
> {{mapGlyphsToChars}} takes each glyph's character from 
> {{findCharacterFromGlyphIndex}}, whose contract is _"if more than one 
> correspondence exists, then the first one is returned (ordered by 
> bfentries\[])"_. A CJK font maps a radical and the ideograph it is the 
> radical of to one glyph — read out of SourceHanSansCN-Medium's cmap, U+2F63 
> and U+751F are both glyph 18742, U+2F45 and U+65B9 are both 14819, U+2F08 and 
> U+4EBA are both 8966. The radical is the lower code point, so the reverse 
> lookup hands the layout the radical, and that is what reaches the PDF's 
> ToUnicode.
> Nothing about the substitution itself is at fault: the round trip through 
> characters loses the identity of the character even where GSUB left the glyph 
> untouched.
> h4. Fix
> {{mapGlyphsToChars}} should prefer the character the glyph came from — the 
> {{GlyphSequence}} already carries it in the glyph's {{CharAssociation}} — 
> wherever the substitution left the glyph alone, i.e. the association covers 
> exactly one character and the font's character map maps that character to 
> this same glyph:
> {noformat}
> int gi = gs.getGlyph(i);
> int cc = findUnsubstitutedCharacter(gs, i, ca, nc, gi);
> if (cc == 0) {
>     cc = findCharacterFromGlyphIndex(gi);
> }
> {noformat}
> Anything the substitution did produce — a ligature, or a glyph with no 
> character of its own — still goes through {{findCharacterFromGlyphIndex}} 
> exactly as before, and a supplementary plane character still comes back as 
> its surrogate pair.
> h4. Test
> {{fop-core/src/test/java/org/apache/fop/fonts/MultiByteFontTestCase.java}}, 
> four cases against a hand-built character map holding the three 
> radical/ideograph pairs above, so no font need be installed:
> * an ideograph sharing its glyph with a radical comes back as the ideograph 
> (fails without the fix: {{expected:<\[生方人]牋> but was:<\[⽣⽅⼈]牋>}}),
> * a radical that really was written stays a radical,
> * a glyph the substitution did produce is still mapped through the character 
> map,
> * a supplementary plane character still comes back as its surrogate pair.
> fop-core's full suite passes with the fix: 3648 tests, 0 failures, 4 skipped.
> h4. Patch
> In the PR: one commit, two files, 
> {{fop-core/src/main/java/org/apache/fop/fonts/MultiByteFont.java}} (+40 −1) 
> and the new test (+140).



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to