Ramin Gharib created FLINK-40871:
------------------------------------

             Summary: Avoid the UTF-8 round trip when casting strings and 
decimals to VARIANT
                 Key: FLINK-40871
                 URL: https://issues.apache.org/jira/browse/FLINK-40871
             Project: Flink
          Issue Type: Sub-task
          Components: Table SQL / Runtime
            Reporter: Ramin Gharib
            Assignee: Ramin Gharib


FLINK-40825 added \{{CAST}} from primitive types to VARIANT. Two of its runtime 
helpers do more work per record than needed.

h3. Problem

*Strings.* \{{VariantCastUtils#fromString}} calls \{{StringData#toString()}}, 
which decodes the UTF-8 bytes into a Java String. 
\{{BinaryVariantInternalBuilder#appendString}} then encodes the String back to 
UTF-8, and \{{build()}} copies the result once more. The bytes that land in the 
VARIANT are, for valid input, the bytes we started with.

*Decimals.* \{{VariantCastUtils#fromDecimal}} calls 
\{{DecimalData#toBigDecimal()}}. A compact decimal, with a precision of 18 or 
less, already holds an unscaled long. The BigDecimal is only built for 
\{{appendDecimal}} to take it apart again.

h3. The catch: UTF-8 validity

The Variant spec requires string values to be valid UTF-8. A 
\{{BinaryStringData}} is not guaranteed to hold valid UTF-8. Today the decode 
step hides this: \{{StringUtf8Utils#decodeUTF8}} falls back to \{{new 
String(bytes, UTF_8)}}, which replaces every malformed sequence with U+FFFD. So 
a VARIANT string is always valid UTF-8, if lossy.

Copying \{{BinaryStringData#toBytes()}} straight into the builder would put 
invalid bytes into the VARIANT. Any byte path therefore needs a validation step.

h3. Proposal

* Add an \{{appendString(byte[] utf8)}} overload to 
\{{BinaryVariantInternalBuilder}}, and let \{{appendString(String)}} delegate 
to it.
* In \{{fromString}}, validate the bytes in one pass. Valid bytes go to the new 
overload. Invalid bytes fall back to today's \{{toString()}} path. This keeps 
the result identical to today for every input, and a scan is cheaper than a 
decode plus an encode.
* Add a compact path for decimals that writes the unscaled long directly, 
picking decimal4 or decimal8 the same way \{{appendDecimal}} does today.

Rejecting invalid UTF-8 instead of replacing it is an option too. It would 
change behavior and make every string cast to VARIANT fallible, so it should be 
a separate decision.

h3. Acceptance

* For valid UTF-8, the stored VARIANT is byte-for-byte identical to today's.
* Invalid UTF-8 still stores U+FFFD, exactly as today.
* Compact and non-compact decimals produce the same VARIANT as today, at every 
width boundary.
* A micro benchmark shows the gain for both casts.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to