[
https://issues.apache.org/jira/browse/FLINK-40854?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Timo Walther closed FLINK-40854.
--------------------------------
Fix Version/s: 2.4.0
Release Note: PARSE_JSON and TRY_PARSE_JSON fail on non UTF-8 input. Use
make MAKE_VALID_UTF8 if necessary.
Resolution: Fixed
Fixed in master: f18d273d48bb5d4d75e3827177a5437ebba0a3b9
> Parse UTF-8 bytes directly in PARSE_JSON and TRY_PARSE_JSON
> -----------------------------------------------------------
>
> Key: FLINK-40854
> URL: https://issues.apache.org/jira/browse/FLINK-40854
> Project: Flink
> Issue Type: Improvement
> Components: Table SQL / Runtime
> Reporter: Ramin Gharib
> Assignee: Ramin Gharib
> Priority: Major
> Labels: pull-request-available
> Fix For: 2.4.0
>
>
> *Problem*
> {\{PARSE_JSON}} and \{{TRY_PARSE_JSON}} call \{{StringData.toString()}} and
> parse the resulting Java \{{String}}. A Flink \{{STRING}} is already stored
> as UTF-8 bytes. So every record is first decoded into a \{{String}}, and
> Jackson then walks its chars.
> {code}
> Current: UTF-8 bytes --decode--> String --Jackson char parser--> VARIANT
> Proposed: UTF-8 bytes ---------------------Jackson byte parser--> VARIANT
> {code}
> *Proposed change*
> * Add \{{BinaryVariantInternalBuilder.parseJson(byte[], boolean)}} and call
> it from both functions with \{{StringData.toBytes()}}.
> * Disable Jackson's \{{JsonFactory.Feature.CHARSET_DETECTION}} on that path.
> With detection on, Jackson guesses UTF-16 or UTF-32 from NUL bytes or a BOM
> in the first four bytes. \{{TRY_PARSE_JSON(CONCAT('1', CHR(0)))}} would then
> return \{{1}} instead of \{{NULL}}.
> *Performance*
> JMH 1.37, JDK 21, Apple M5 Max, 2 forks. Each call wraps a fresh
> \{{BinaryStringData}} around bytes at a non-zero offset in a
> \{{MemorySegment}}, like \{{BinaryRowData.getString()}} does per record.
> "Long text" is a few fields plus one long string value. "Records" is an array
> of order objects with numbers, booleans, nested objects and short strings.
> ||Document||Size||Before||After||Change||
> |Long text, ASCII|1 KB|0.95 µs|0.57 µs|-40%|
> |Long text, ASCII|100 KB|160 µs|63 µs|-61%|
> |Long text, ASCII|1 MB|1.71 ms|0.63 ms|-63%|
> |Long text, non-ASCII|100 KB|194 µs|149 µs|-23%|
> |Long text, non-ASCII|1 MB|2.06 ms|1.54 ms|-25%|
> |Records, non-ASCII|830 B|2.65 µs|2.26 µs|-15%|
> |Records, ASCII|100 KB|383 µs|340 µs|-11%|
> |Records, non-ASCII|100 KB|431 µs|384 µs|-11%|
> |Records, non-ASCII|1 MB|4.17 ms|3.61 ms|-13%|
> Small record-shaped documents of 1 to 10 KB are roughly flat. The gain is
> largest for long string values, where the old decode and copy dominate.
> Allocation per call drops by 22 to 29% for documents of 100 KB and more.
> *Behavior change*
> Valid JSON parses to the same VARIANT as before. Only a \{{STRING}} that
> holds invalid UTF-8 behaves differently. Such strings come from sources that
> wrap bytes without validating them.
> ||Input bytes||Function||Before||After||
> |\{{22 FF 22}}|\{{PARSE_JSON}}|\{{"\uFFFD"}}|error|
> |\{{22 FF 22}}|\{{TRY_PARSE_JSON}}|\{{"\uFFFD"}}|\{{NULL}}|
> |\{{22 C0 AF 22}}|\{{PARSE_JSON}}|\{{"\uFFFD\uFFFD"}}|\{{"/"}}|
> The first two rows match FLINK-39623, which made \{{CAST(BYTES AS STRING)}}
> fail on invalid UTF-8 by default. The last row comes from Jackson decoding
> overlong UTF-8 sequences instead of rejecting them. Rejecting those too would
> need a separate UTF-8 validation pass. Open question whether that is worth
> the cost.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)