David Mollitor created SPARK-59389:
--------------------------------------
Summary: Avoid per-character UTF8String allocation in
trimLeft/trimRight with a trim set
Key: SPARK-59389
URL: https://issues.apache.org/jira/browse/SPARK-59389
Project: Spark
Issue Type: Improvement
Components: Spark Core
Affects Versions: 4.1.0
Reporter: David Mollitor
h2. Summary
The trim overloads that take a set of trim characters --
{{trimLeft(UTF8String)}} and
{{trimRight(UTF8String)}}, with {{trim(UTF8String)}} delegating to both --
allocate a throwaway {{UTF8String}} for every source character they examine.
Each iteration calls {{copyUTF8String(...)}} (a {{new byte[]}} + {{copyMemory}}
+ {{fromBytes}}) only to pass that one character to
{{trimString.find(searchChar, 0)}}. Trimming k leading/trailing characters does
O(k) heap allocations.
The character under test is already a byte range in this string's backing
memory, so the
membership check can read it directly. This factors {{find}} to accept a raw
{{(base, offset, numBytes)}} needle and calls it on the source character's
bytes in place, so the O(k) allocations become O(1). Performance-only; behavior
is identical.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]