Ramin Gharib created FLINK-40871:
------------------------------------
Summary: Avoid the UTF-8 round trip when casting strings and
decimals to VARIANT
Key: FLINK-40871
URL: https://issues.apache.org/jira/browse/FLINK-40871
Project: Flink
Issue Type: Sub-task
Components: Table SQL / Runtime
Reporter: Ramin Gharib
Assignee: Ramin Gharib
FLINK-40825 added \{{CAST}} from primitive types to VARIANT. Two of its runtime
helpers do more work per record than needed.
h3. Problem
*Strings.* \{{VariantCastUtils#fromString}} calls \{{StringData#toString()}},
which decodes the UTF-8 bytes into a Java String.
\{{BinaryVariantInternalBuilder#appendString}} then encodes the String back to
UTF-8, and \{{build()}} copies the result once more. The bytes that land in the
VARIANT are, for valid input, the bytes we started with.
*Decimals.* \{{VariantCastUtils#fromDecimal}} calls
\{{DecimalData#toBigDecimal()}}. A compact decimal, with a precision of 18 or
less, already holds an unscaled long. The BigDecimal is only built for
\{{appendDecimal}} to take it apart again.
h3. The catch: UTF-8 validity
The Variant spec requires string values to be valid UTF-8. A
\{{BinaryStringData}} is not guaranteed to hold valid UTF-8. Today the decode
step hides this: \{{StringUtf8Utils#decodeUTF8}} falls back to \{{new
String(bytes, UTF_8)}}, which replaces every malformed sequence with U+FFFD. So
a VARIANT string is always valid UTF-8, if lossy.
Copying \{{BinaryStringData#toBytes()}} straight into the builder would put
invalid bytes into the VARIANT. Any byte path therefore needs a validation step.
h3. Proposal
* Add an \{{appendString(byte[] utf8)}} overload to
\{{BinaryVariantInternalBuilder}}, and let \{{appendString(String)}} delegate
to it.
* In \{{fromString}}, validate the bytes in one pass. Valid bytes go to the new
overload. Invalid bytes fall back to today's \{{toString()}} path. This keeps
the result identical to today for every input, and a scan is cheaper than a
decode plus an encode.
* Add a compact path for decimals that writes the unscaled long directly,
picking decimal4 or decimal8 the same way \{{appendDecimal}} does today.
Rejecting invalid UTF-8 instead of replacing it is an option too. It would
change behavior and make every string cast to VARIANT fallible, so it should be
a separate decision.
h3. Acceptance
* For valid UTF-8, the stored VARIANT is byte-for-byte identical to today's.
* Invalid UTF-8 still stores U+FFFD, exactly as today.
* Compact and non-compact decimals produce the same VARIANT as today, at every
width boundary.
* A micro benchmark shows the gain for both casts.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)