[ 
https://issues.apache.org/jira/browse/FOP-3345?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Jason Harrop updated FOP-3345:
------------------------------
    Description: 
MultiByteFont.mapGlyphsToChars maps the substituted glyph sequence back to 
characters one character per glyph. For a glyph substitution left alone it 
finds the originating character. For a glyph substitution produced, whose glyph 
index no character maps to, it falls through to createPrivateUseMapping, which 
mints a code point from U+E000 upward. The painter later hands that code point 
to CIDSubset.mapCodePoint, and CIDSubset.getChars is what PDFToUnicodeCMap is 
built from. So the private-use code point minted to give the glyph an identity 
becomes its published meaning.

That is every ligature (fi, fl, ffi, ft, ti and so on wherever the font has no 
cmap entry for the ligature glyph), and every single substitution whose output 
glyph has no cmap entry of its own: Arabic initial, medial and final forms in a 
font without presentation-form cmap entries, small capitals, old-style figures, 
contextual alternates, and each glyph of a ccmp decomposition.

Measured with the 2.11 command line, read back with pdftotext:

{noformat}
Carlito, script="latn", "office affluent fifty flow ti fi"
    -> "office affluent fiy flow  fi"
       (ft and ti minted; fi, fl, ffi, ffl publish U+FB01..FB04 because 
Carlito's
       cmap happens to map those forms)
Noto Sans Arabic, script="arab", "السلام عليكم"
    -> the medial yeh extracts as U+E001 U+E000 (ccmp split it into 
uni066E.medi.wide
       and twodotshorizontalbelowar), the other letters as presentation forms
{noformat}

A docx4j corpus render of an Arabic probe through 2.11 counted 105 private-use 
characters against 308 base letters and 441 presentation forms.

The page looks right, since the glyph and its advance are correct; the defect 
shows only when someone selects, searches or reads the text.

h3. Fix

The information is already there: the GlyphSequence carries a CharAssociation 
per glyph naming the characters that produced it, and 
findUnsubstitutedCharacter reads it and discards it for exactly the substituted 
case.

# MultiByteFont.mapGlyphsToChars records, per glyph index, the characters of 
the association whenever the glyph was substituted (findUnsubstitutedCharacter 
returned nothing). A disjoint association (a ligature with ignored marks 
between its components) contributes its sub-intervals only. The character 
returned for layout does not change: it is still the glyph's identity, which 
findGlyphIndex maps back to the glyph.
# CIDSet gains getUnicodeSequences(): one String per selector, the recorded 
characters where there are any and the code point otherwise. CIDSubset and 
CIDFull implement it.
# PDFFactory builds the ToUnicode CMap from that instead of getChars().
# PDFToUnicodeCMap takes one destination String per selector. Only a 
destination of one code point may join a bfrange; a multi-character destination 
is written as a bfchar with a string, which the spec allows.

Two cases a ToUnicode entry cannot express, and how the fix leaves them:

* Several glyphs for one character (a ccmp decomposition replicates the 
association onto each output glyph). The first glyph publishes the character; 
the second and later keep the private-use code point. An empty destination was 
tried and rejected: poppler reads it correctly, but mupdf prints U+FFFD, pdf.js 
a space, and PDFium the raw selector as a control character. Publishing the 
character twice makes poppler and mupdf insert a space. ActualText per cluster 
is the mechanism for this case and is a content-stream change outside this 
issue.
* One glyph seen with two different meanings (a shared dotless base in a 
decomposing Arabic font). It publishes neither, keeping the private-use code 
point, rather than the wrong letter.

After the fix, with the drawn content byte for byte unchanged:

{noformat}
Carlito         -> "office affluent fifty flow ti fi"   (pdftotext, mupdf, 
pdf.js, PDFium)
Noto Sans Arabic -> "السلام عليكم"                 (the dots glyph is the one 
left)
{noformat}

Tests: PDFToUnicodeCMapTestCase and MultiByteFontTestCase extended; the 
writer's existing output is pinned by ToUnicodeCharacterisationTestCase before 
the change.

Related: the same rewrite of PDFToUnicodeCMap fixes the selector drift after a 
surrogate pair, reported separately (FOP-3346), because the positional char\[] 
that caused it is what had to go.


  was:
MultiByteFont.mapGlyphsToChars maps the substituted glyph sequence back to 
characters one character per glyph. For a glyph substitution left alone it 
finds the originating character. For a glyph substitution produced, whose glyph 
index no character maps to, it falls through to createPrivateUseMapping, which 
mints a code point from U+E000 upward. The painter later hands that code point 
to CIDSubset.mapCodePoint, and CIDSubset.getChars is what PDFToUnicodeCMap is 
built from. So the private-use code point minted to give the glyph an identity 
becomes its published meaning.

That is every ligature (fi, fl, ffi, ft, ti and so on wherever the font has no 
cmap entry for the ligature glyph), and every single substitution whose output 
glyph has no cmap entry of its own: Arabic initial, medial and final forms in a 
font without presentation-form cmap entries, small capitals, old-style figures, 
contextual alternates, and each glyph of a ccmp decomposition.

Measured with the 2.11 command line, read back with pdftotext:

{noformat}
Carlito, script="latn", "office affluent fifty flow ti fi"
    -> "office affluent fiy flow  fi"
       (ft and ti minted; fi, fl, ffi, ffl publish U+FB01..FB04 because 
Carlito's
       cmap happens to map those forms)
Noto Sans Arabic, script="arab", "السلام عليكم"
    -> the medial yeh extracts as U+E001 U+E000 (ccmp split it into 
uni066E.medi.wide
       and twodotshorizontalbelowar), the other letters as presentation forms
{noformat}

A docx4j corpus render of an Arabic probe through 2.11 counted 105 private-use 
characters against 308 base letters and 441 presentation forms.

The page looks right, since the glyph and its advance are correct; the defect 
shows only when someone selects, searches or reads the text.

h3. Fix

The information is already there: the GlyphSequence carries a CharAssociation 
per glyph naming the characters that produced it, and 
findUnsubstitutedCharacter reads it and discards it for exactly the substituted 
case.

# MultiByteFont.mapGlyphsToChars records, per glyph index, the characters of 
the association whenever the glyph was substituted (findUnsubstitutedCharacter 
returned nothing). A disjoint association (a ligature with ignored marks 
between its components) contributes its sub-intervals only. The character 
returned for layout does not change: it is still the glyph's identity, which 
findGlyphIndex maps back to the glyph.
# CIDSet gains getUnicodeSequences(): one String per selector, the recorded 
characters where there are any and the code point otherwise. CIDSubset and 
CIDFull implement it.
# PDFFactory builds the ToUnicode CMap from that instead of getChars().
# PDFToUnicodeCMap takes one destination String per selector. Only a 
destination of one code point may join a bfrange; a multi-character destination 
is written as a bfchar with a string, which the spec allows.

Two cases a ToUnicode entry cannot express, and how the fix leaves them:

* Several glyphs for one character (a ccmp decomposition replicates the 
association onto each output glyph). The first glyph publishes the character; 
the second and later keep the private-use code point. An empty destination was 
tried and rejected: poppler reads it correctly, but mupdf prints U+FFFD, pdf.js 
a space, and PDFium the raw selector as a control character. Publishing the 
character twice makes poppler and mupdf insert a space. ActualText per cluster 
is the mechanism for this case and is a content-stream change outside this 
issue.
* One glyph seen with two different meanings (a shared dotless base in a 
decomposing Arabic font). It publishes neither, keeping the private-use code 
point, rather than the wrong letter.

After the fix, with the drawn content byte for byte unchanged:

{noformat}
Carlito         -> "office affluent fifty flow ti fi"   (pdftotext, mupdf, 
pdf.js, PDFium)
Noto Sans Arabic -> "السلام عليكم"                 (the dots glyph is the one 
left)
{noformat}

Tests: PDFToUnicodeCMapTestCase and MultiByteFontTestCase extended; the 
writer's existing output is pinned by ToUnicodeCharacterisationTestCase before 
the change.

Related: the same rewrite of PDFToUnicodeCMap fixes the selector drift after a 
surrogate pair, reported separately (a companion issue filed alongside this 
one), because the positional char\[] that caused it is what had to go.



> ToUnicode publishes a private-use code point for any glyph substitution 
> produced, so ligatures and Arabic contextual forms are not searchable, 
> copyable or readable aloud
> -------------------------------------------------------------------------------------------------------------------------------------------------------------------------
>
>                 Key: FOP-3345
>                 URL: https://issues.apache.org/jira/browse/FOP-3345
>             Project: FOP
>          Issue Type: Bug
>          Components: font/opentype, renderer/pdf
>    Affects Versions: 2.11
>            Reporter: Jason Harrop
>            Priority: Minor
>
> MultiByteFont.mapGlyphsToChars maps the substituted glyph sequence back to 
> characters one character per glyph. For a glyph substitution left alone it 
> finds the originating character. For a glyph substitution produced, whose 
> glyph index no character maps to, it falls through to 
> createPrivateUseMapping, which mints a code point from U+E000 upward. The 
> painter later hands that code point to CIDSubset.mapCodePoint, and 
> CIDSubset.getChars is what PDFToUnicodeCMap is built from. So the private-use 
> code point minted to give the glyph an identity becomes its published meaning.
> That is every ligature (fi, fl, ffi, ft, ti and so on wherever the font has 
> no cmap entry for the ligature glyph), and every single substitution whose 
> output glyph has no cmap entry of its own: Arabic initial, medial and final 
> forms in a font without presentation-form cmap entries, small capitals, 
> old-style figures, contextual alternates, and each glyph of a ccmp 
> decomposition.
> Measured with the 2.11 command line, read back with pdftotext:
> {noformat}
> Carlito, script="latn", "office affluent fifty flow ti fi"
>     -> "office affluent fiy flow  fi"
>        (ft and ti minted; fi, fl, ffi, ffl publish U+FB01..FB04 because 
> Carlito's
>        cmap happens to map those forms)
> Noto Sans Arabic, script="arab", "السلام عليكم"
>     -> the medial yeh extracts as U+E001 U+E000 (ccmp split it into 
> uni066E.medi.wide
>        and twodotshorizontalbelowar), the other letters as presentation forms
> {noformat}
> A docx4j corpus render of an Arabic probe through 2.11 counted 105 
> private-use characters against 308 base letters and 441 presentation forms.
> The page looks right, since the glyph and its advance are correct; the defect 
> shows only when someone selects, searches or reads the text.
> h3. Fix
> The information is already there: the GlyphSequence carries a CharAssociation 
> per glyph naming the characters that produced it, and 
> findUnsubstitutedCharacter reads it and discards it for exactly the 
> substituted case.
> # MultiByteFont.mapGlyphsToChars records, per glyph index, the characters of 
> the association whenever the glyph was substituted 
> (findUnsubstitutedCharacter returned nothing). A disjoint association (a 
> ligature with ignored marks between its components) contributes its 
> sub-intervals only. The character returned for layout does not change: it is 
> still the glyph's identity, which findGlyphIndex maps back to the glyph.
> # CIDSet gains getUnicodeSequences(): one String per selector, the recorded 
> characters where there are any and the code point otherwise. CIDSubset and 
> CIDFull implement it.
> # PDFFactory builds the ToUnicode CMap from that instead of getChars().
> # PDFToUnicodeCMap takes one destination String per selector. Only a 
> destination of one code point may join a bfrange; a multi-character 
> destination is written as a bfchar with a string, which the spec allows.
> Two cases a ToUnicode entry cannot express, and how the fix leaves them:
> * Several glyphs for one character (a ccmp decomposition replicates the 
> association onto each output glyph). The first glyph publishes the character; 
> the second and later keep the private-use code point. An empty destination 
> was tried and rejected: poppler reads it correctly, but mupdf prints U+FFFD, 
> pdf.js a space, and PDFium the raw selector as a control character. 
> Publishing the character twice makes poppler and mupdf insert a space. 
> ActualText per cluster is the mechanism for this case and is a content-stream 
> change outside this issue.
> * One glyph seen with two different meanings (a shared dotless base in a 
> decomposing Arabic font). It publishes neither, keeping the private-use code 
> point, rather than the wrong letter.
> After the fix, with the drawn content byte for byte unchanged:
> {noformat}
> Carlito         -> "office affluent fifty flow ti fi"   (pdftotext, mupdf, 
> pdf.js, PDFium)
> Noto Sans Arabic -> "السلام عليكم"                 (the dots glyph is the 
> one left)
> {noformat}
> Tests: PDFToUnicodeCMapTestCase and MultiByteFontTestCase extended; the 
> writer's existing output is pinned by ToUnicodeCharacterisationTestCase 
> before the change.
> Related: the same rewrite of PDFToUnicodeCMap fixes the selector drift after 
> a surrogate pair, reported separately (FOP-3346), because the positional 
> char\[] that caused it is what had to go.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to