[
https://issues.apache.org/jira/browse/SPARK-59378?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
David Mollitor resolved SPARK-59378.
------------------------------------
Resolution: Won't Fix
Caching the full-ASCII flag as a byproduct of {{UTF8String.numChars()}} (this
issue) and the ASCII fast-path for the char/byte position-lookup methods
(SPARK-59379) only deliver value together: on its own, this change writes a
flag that almost nothing on master reads, and SPARK-59379 reads a flag that
almost nothing on master writes. Per reviewer feedback on the pull requests,
the two were merged into a single change under SPARK-59379), and this issue's
PR was closed.
> Cache the full-ASCII flag as a byproduct of UTF8String.numChars()
> -----------------------------------------------------------------
>
> Key: SPARK-59378
> URL: https://issues.apache.org/jira/browse/SPARK-59378
> Project: Spark
> Issue Type: Improvement
> Components: Spark Core
> Affects Versions: 4.1.0
> Reporter: David Mollitor
> Priority: Minor
> Labels: pull-request-available
>
> h2. Summary
> {{UTF8String.numChars()}} scans every lead byte of the string to count code
> points.
> Determining whether a string is full ASCII ({{{}isFullAscii(){}}} /
> {{{}getIsFullAscii(){}}})
> requires the same kind of full scan. When both are needed for the same
> string, the
> string is scanned twice.
> This change computes the full-ASCII flag as a byproduct of the code-point
> count in
> {{getNumChars()}} and caches it in the existing {{isFullAscii}} field, so a
> subsequent
> {{isFullAscii()}} call is answered from cache instead of triggering a second
> scan.
> h2. Details
> * {{getNumChars()}} now reads each lead byte once into a local and uses it
> both for the width step ({{{}numBytesForFirstByte{}}}) and the ASCII check.
> * The ASCII predicate is {{b >= 0}} per lead byte, which is exactly
> equivalent to
> {{{}getIsFullAscii(){}}}'s {{getByte(i) < 0}} check, including for invalid
> UTF-8. Any
> multi-byte lead ({{{}0xC2{}}} to {{{}0xF4{}}}) or invalid byte ({{{}0x80{}}}
> to {{{}0xC1{}}}, {{0xF5}}
> to {{{}0xFF{}}}) is negative and trips the flag; continuation bytes of a
> valid sequence are
> skipped, but their lead byte already tripped it.
> * The flag is only set when still {{{}UNKNOWN{}}}, preserving any value
> already computed by {{{}getIsFullAscii(){}}}. Both compute the same
> predicate, so they never disagree.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]