[ 
https://issues.apache.org/jira/browse/FLINK-40854?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

ASF GitHub Bot updated FLINK-40854:
-----------------------------------
    Labels: pull-request-available  (was: )

> Parse UTF-8 bytes directly in PARSE_JSON and TRY_PARSE_JSON
> -----------------------------------------------------------
>
>                 Key: FLINK-40854
>                 URL: https://issues.apache.org/jira/browse/FLINK-40854
>             Project: Flink
>          Issue Type: Improvement
>          Components: Table SQL / Runtime
>            Reporter: Ramin Gharib
>            Assignee: Ramin Gharib
>            Priority: Major
>              Labels: pull-request-available
>
> *Problem*
> {\{PARSE_JSON}} and \{{TRY_PARSE_JSON}} call \{{StringData.toString()}} and 
> parse the resulting Java \{{String}}. A Flink \{{STRING}} is already stored 
> as UTF-8 bytes. So every record is first decoded into a \{{String}}, and 
> Jackson then walks its chars.
> {code}
> Current:  UTF-8 bytes --decode--> String --Jackson char parser--> VARIANT
> Proposed: UTF-8 bytes ---------------------Jackson byte parser--> VARIANT
> {code}
> *Proposed change*
> * Add \{{BinaryVariantInternalBuilder.parseJson(byte[], boolean)}} and call 
> it from both functions with \{{StringData.toBytes()}}.
> * Disable Jackson's \{{JsonFactory.Feature.CHARSET_DETECTION}} on that path. 
> With detection on, Jackson guesses UTF-16 or UTF-32 from NUL bytes or a BOM 
> in the first four bytes. \{{TRY_PARSE_JSON(CONCAT('1', CHR(0)))}} would then 
> return \{{1}} instead of \{{NULL}}.
> *Performance*
> JMH 1.37, JDK 21, Apple M5 Max, 2 forks. Each call wraps a fresh 
> \{{BinaryStringData}} around bytes at a non-zero offset in a 
> \{{MemorySegment}}, like \{{BinaryRowData.getString()}} does per record. 
> "Long text" is a few fields plus one long string value. "Records" is an array 
> of order objects with numbers, booleans, nested objects and short strings.
> ||Document||Size||Before||After||Change||
> |Long text, ASCII|1 KB|0.95 µs|0.57 µs|-40%|
> |Long text, ASCII|100 KB|160 µs|63 µs|-61%|
> |Long text, ASCII|1 MB|1.71 ms|0.63 ms|-63%|
> |Long text, non-ASCII|100 KB|194 µs|149 µs|-23%|
> |Long text, non-ASCII|1 MB|2.06 ms|1.54 ms|-25%|
> |Records, non-ASCII|830 B|2.65 µs|2.26 µs|-15%|
> |Records, ASCII|100 KB|383 µs|340 µs|-11%|
> |Records, non-ASCII|100 KB|431 µs|384 µs|-11%|
> |Records, non-ASCII|1 MB|4.17 ms|3.61 ms|-13%|
> Small record-shaped documents of 1 to 10 KB are roughly flat. The gain is 
> largest for long string values, where the old decode and copy dominate. 
> Allocation per call drops by 22 to 29% for documents of 100 KB and more.
> *Behavior change*
> Valid JSON parses to the same VARIANT as before. Only a \{{STRING}} that 
> holds invalid UTF-8 behaves differently. Such strings come from sources that 
> wrap bytes without validating them.
> ||Input bytes||Function||Before||After||
> |\{{22 FF 22}}|\{{PARSE_JSON}}|\{{"\uFFFD"}}|error|
> |\{{22 FF 22}}|\{{TRY_PARSE_JSON}}|\{{"\uFFFD"}}|\{{NULL}}|
> |\{{22 C0 AF 22}}|\{{PARSE_JSON}}|\{{"\uFFFD\uFFFD"}}|\{{"/"}}|
> The first two rows match FLINK-39623, which made \{{CAST(BYTES AS STRING)}} 
> fail on invalid UTF-8 by default. The last row comes from Jackson decoding 
> overlong UTF-8 sequences instead of rejecting them. Rejecting those too would 
> need a separate UTF-8 validation pass. Open question whether that is worth 
> the cost.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to