Ramin Gharib created FLINK-40854:
------------------------------------
Summary: Parse UTF-8 bytes directly in PARSE_JSON and
TRY_PARSE_JSON
Key: FLINK-40854
URL: https://issues.apache.org/jira/browse/FLINK-40854
Project: Flink
Issue Type: Improvement
Components: Table SQL / Runtime
Reporter: Ramin Gharib
Assignee: Ramin Gharib
*Problem*
{\{PARSE_JSON}} and \{{TRY_PARSE_JSON}} call \{{StringData.toString()}} and
parse the resulting Java \{{String}}. A Flink \{{STRING}} is already stored as
UTF-8 bytes. So every record is first decoded into a \{{String}}, and Jackson
then walks its chars.
{code}
Current: UTF-8 bytes --decode--> String --Jackson char parser--> VARIANT
Proposed: UTF-8 bytes ---------------------Jackson byte parser--> VARIANT
{code}
*Proposed change*
* Add \{{BinaryVariantInternalBuilder.parseJson(byte[], boolean)}} and call it
from both functions with \{{StringData.toBytes()}}.
* Disable Jackson's \{{JsonFactory.Feature.CHARSET_DETECTION}} on that path.
With detection on, Jackson guesses UTF-16 or UTF-32 from NUL bytes or a BOM in
the first four bytes. \{{TRY_PARSE_JSON(CONCAT('1', CHR(0)))}} would then
return \{{1}} instead of \{{NULL}}.
*Performance*
JMH 1.37, JDK 21, Apple M5 Max, 2 forks. Each call wraps a fresh
\{{BinaryStringData}} around bytes at a non-zero offset in a
\{{MemorySegment}}, like \{{BinaryRowData.getString()}} does per record. "Long
text" is a few fields plus one long string value. "Records" is an array of
order objects with numbers, booleans, nested objects and short strings.
||Document||Size||Before||After||Change||
|Long text, ASCII|1 KB|0.95 µs|0.57 µs|-40%|
|Long text, ASCII|100 KB|160 µs|63 µs|-61%|
|Long text, ASCII|1 MB|1.71 ms|0.63 ms|-63%|
|Long text, non-ASCII|100 KB|194 µs|149 µs|-23%|
|Long text, non-ASCII|1 MB|2.06 ms|1.54 ms|-25%|
|Records, non-ASCII|830 B|2.65 µs|2.26 µs|-15%|
|Records, ASCII|100 KB|383 µs|340 µs|-11%|
|Records, non-ASCII|100 KB|431 µs|384 µs|-11%|
|Records, non-ASCII|1 MB|4.17 ms|3.61 ms|-13%|
Small record-shaped documents of 1 to 10 KB are roughly flat. The gain is
largest for long string values, where the old decode and copy dominate.
Allocation per call drops by 22 to 29% for documents of 100 KB and more.
*Behavior change*
Valid JSON parses to the same VARIANT as before. Only a \{{STRING}} that holds
invalid UTF-8 behaves differently. Such strings come from sources that wrap
bytes without validating them.
||Input bytes||Function||Before||After||
|\{{22 FF 22}}|\{{PARSE_JSON}}|\{{"\uFFFD"}}|error|
|\{{22 FF 22}}|\{{TRY_PARSE_JSON}}|\{{"\uFFFD"}}|\{{NULL}}|
|\{{22 C0 AF 22}}|\{{PARSE_JSON}}|\{{"\uFFFD\uFFFD"}}|\{{"/"}}|
The first two rows match FLINK-39623, which made \{{CAST(BYTES AS STRING)}}
fail on invalid UTF-8 by default. The last row comes from Jackson decoding
overlong UTF-8 sequences instead of rejecting them. Rejecting those too would
need a separate UTF-8 validation pass. Open question whether that is worth the
cost.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)