[
https://issues.apache.org/jira/browse/FLINK-40871?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated FLINK-40871:
-----------------------------------
Labels: pull-request-available (was: )
> Avoid the UTF-8 round trip when casting strings and decimals to VARIANT
> -----------------------------------------------------------------------
>
> Key: FLINK-40871
> URL: https://issues.apache.org/jira/browse/FLINK-40871
> Project: Flink
> Issue Type: Sub-task
> Components: Table SQL / Runtime
> Reporter: Ramin Gharib
> Assignee: Ramin Gharib
> Priority: Major
> Labels: pull-request-available
>
> FLINK-40825 added \{{CAST}} from primitive types to VARIANT. Two of its
> runtime helpers do more work per record than needed.
> h3. Problem
> *Strings.* \{{VariantCastUtils#fromString}} calls \{{StringData#toString()}},
> which decodes the UTF-8 bytes into a Java String.
> \{{BinaryVariantInternalBuilder#appendString}} then encodes the String back
> to UTF-8, and \{{build()}} copies the result once more. The bytes that land
> in the VARIANT are, for valid input, the bytes we started with.
> *Decimals.* \{{VariantCastUtils#fromDecimal}} calls
> \{{DecimalData#toBigDecimal()}}. A compact decimal, with a precision of 18 or
> less, already holds an unscaled long. The BigDecimal is only built for
> \{{appendDecimal}} to take it apart again.
> h3. The catch: UTF-8 validity
> The Variant spec requires string values to be valid UTF-8. A
> \{{BinaryStringData}} is not guaranteed to hold valid UTF-8. Today the decode
> step hides this: \{{StringUtf8Utils#decodeUTF8}} falls back to \{{new
> String(bytes, UTF_8)}}, which replaces every malformed sequence with U+FFFD.
> So a VARIANT string is always valid UTF-8, if lossy.
> Copying \{{BinaryStringData#toBytes()}} straight into the builder would put
> invalid bytes into the VARIANT. Any byte path therefore needs a validation
> step.
> h3. Proposal
> * Add an \{{appendString(byte[] utf8)}} overload to
> \{{BinaryVariantInternalBuilder}}, and let \{{appendString(String)}} delegate
> to it.
> * In \{{fromString}}, validate the bytes in one pass. Valid bytes go to the
> new overload. Invalid bytes fall back to today's \{{toString()}} path. This
> keeps the result identical to today for every input, and a scan is cheaper
> than a decode plus an encode.
> * Add a compact path for decimals that writes the unscaled long directly,
> picking decimal4 or decimal8 the same way \{{appendDecimal}} does today.
> Rejecting invalid UTF-8 instead of replacing it is an option too. It would
> change behavior and make every string cast to VARIANT fallible, so it should
> be a separate decision.
> h3. Acceptance
> * For valid UTF-8, the stored VARIANT is byte-for-byte identical to today's.
> * Invalid UTF-8 still stores U+FFFD, exactly as today.
> * Compact and non-compact decimals produce the same VARIANT as today, at
> every width boundary.
> * A micro benchmark shows the gain for both casts.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)