Ramin Gharib created FLINK-40854:
------------------------------------

             Summary: Parse UTF-8 bytes directly in PARSE_JSON and 
TRY_PARSE_JSON
                 Key: FLINK-40854
                 URL: https://issues.apache.org/jira/browse/FLINK-40854
             Project: Flink
          Issue Type: Improvement
          Components: Table SQL / Runtime
            Reporter: Ramin Gharib
            Assignee: Ramin Gharib


*Problem*

{\{PARSE_JSON}} and \{{TRY_PARSE_JSON}} call \{{StringData.toString()}} and 
parse the resulting Java \{{String}}. A Flink \{{STRING}} is already stored as 
UTF-8 bytes. So every record is first decoded into a \{{String}}, and Jackson 
then walks its chars.

{code}
Current:  UTF-8 bytes --decode--> String --Jackson char parser--> VARIANT
Proposed: UTF-8 bytes ---------------------Jackson byte parser--> VARIANT
{code}

*Proposed change*

* Add \{{BinaryVariantInternalBuilder.parseJson(byte[], boolean)}} and call it 
from both functions with \{{StringData.toBytes()}}.
* Disable Jackson's \{{JsonFactory.Feature.CHARSET_DETECTION}} on that path. 
With detection on, Jackson guesses UTF-16 or UTF-32 from NUL bytes or a BOM in 
the first four bytes. \{{TRY_PARSE_JSON(CONCAT('1', CHR(0)))}} would then 
return \{{1}} instead of \{{NULL}}.

*Performance*

JMH 1.37, JDK 21, Apple M5 Max, 2 forks. Each call wraps a fresh 
\{{BinaryStringData}} around bytes at a non-zero offset in a 
\{{MemorySegment}}, like \{{BinaryRowData.getString()}} does per record. "Long 
text" is a few fields plus one long string value. "Records" is an array of 
order objects with numbers, booleans, nested objects and short strings.

||Document||Size||Before||After||Change||
|Long text, ASCII|1 KB|0.95 µs|0.57 µs|-40%|
|Long text, ASCII|100 KB|160 µs|63 µs|-61%|
|Long text, ASCII|1 MB|1.71 ms|0.63 ms|-63%|
|Long text, non-ASCII|100 KB|194 µs|149 µs|-23%|
|Long text, non-ASCII|1 MB|2.06 ms|1.54 ms|-25%|
|Records, non-ASCII|830 B|2.65 µs|2.26 µs|-15%|
|Records, ASCII|100 KB|383 µs|340 µs|-11%|
|Records, non-ASCII|100 KB|431 µs|384 µs|-11%|
|Records, non-ASCII|1 MB|4.17 ms|3.61 ms|-13%|

Small record-shaped documents of 1 to 10 KB are roughly flat. The gain is 
largest for long string values, where the old decode and copy dominate. 
Allocation per call drops by 22 to 29% for documents of 100 KB and more.

*Behavior change*

Valid JSON parses to the same VARIANT as before. Only a \{{STRING}} that holds 
invalid UTF-8 behaves differently. Such strings come from sources that wrap 
bytes without validating them.

||Input bytes||Function||Before||After||
|\{{22 FF 22}}|\{{PARSE_JSON}}|\{{"\uFFFD"}}|error|
|\{{22 FF 22}}|\{{TRY_PARSE_JSON}}|\{{"\uFFFD"}}|\{{NULL}}|
|\{{22 C0 AF 22}}|\{{PARSE_JSON}}|\{{"\uFFFD\uFFFD"}}|\{{"/"}}|

The first two rows match FLINK-39623, which made \{{CAST(BYTES AS STRING)}} 
fail on invalid UTF-8 by default. The last row comes from Jackson decoding 
overlong UTF-8 sequences instead of rejecting them. Rejecting those too would 
need a separate UTF-8 validation pass. Open question whether that is worth the 
cost.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to