[ 
https://issues.apache.org/jira/browse/FLINK-40305?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Ramin Gharib updated FLINK-40305:
---------------------------------
    Description: 
{{BinaryVariantUtil}} decodes variant strings and object field names without 
passing a charset, so both fall back to the JVM default:
 
{code:java}
// getString
return new String(value, start, length);
// getMetadataKey
return new String(metadata, stringStart + offset, nextOffset - offset);
{code}
 
Both are written as UTF-8 by {{{}BinaryVariantInternalBuilder{}}}, which uses 
{{str.getBytes(StandardCharsets.UTF_8)}} for values and 
{{key.getBytes(StandardCharsets.UTF_8)}} for keys, so the write and read paths 
disagree.
 
The default charset is only guaranteed to be UTF-8 from Java 18 onwards (JEP 
400). On Java 11 and 17 it is platform dependent, so a variant holding 
non-ASCII text can decode incorrectly. {{new String(byte[], ...)}} also 
substitutes {{U+FFFD}} for malformed input rather than failing.
 
{{getMetadataKey}} is the more damaging of the two, since it decodes field 
names. A mangled name makes {{getField(name)}} silently return null and 
corrupts {{getFieldNames()}} and {{{}toJson(){}}}.
 
This is a regression from the original port. Spark's {{{}VariantUtil [1]{}}}, 
which this class is derived from, passes the charset in both places:
 
{code:java}
return new String(value, start, length, StandardCharsets.UTF_8);
return new String(metadata, stringStart + offset, nextOffset - offset, 
StandardCharsets.UTF_8);
{code}
h3. Fix

Pass {{StandardCharsets.UTF_8}} in both calls. Add a round-trip test with a 
non-ASCII string value and a non-ASCII field name.

 

 [1] 
https://github.com/apache/spark/blob/master/common/variant/src/main/java/org/apache/spark/types/variant/VariantUtil.java#L617

  was:
{\{BinaryVariantUtil}} decodes variant strings and object field names without 
passing a charset, so both fall back to the JVM default:
 
{code:java}
// getString
return new String(value, start, length);
// getMetadataKey
return new String(metadata, stringStart + offset, nextOffset - offset);
{code}
 
Both are written as UTF-8 by \{{BinaryVariantInternalBuilder}}, which uses 
\{{str.getBytes(StandardCharsets.UTF_8)}} for values and 
\{{key.getBytes(StandardCharsets.UTF_8)}} for keys, so the write and read paths 
disagree.
 
The default charset is only guaranteed to be UTF-8 from Java 18 onwards (JEP 
400). On Java 11 and 17 it is platform dependent, so a variant holding 
non-ASCII text can decode incorrectly. \{{new String(byte[], ...)}} also 
substitutes \{{U+FFFD}} for malformed input rather than failing.
 
{\{getMetadataKey}} is the more damaging of the two, since it decodes field 
names. A mangled name makes \{{getField(name)}} silently return null and 
corrupts \{{getFieldNames()}} and \{{toJson()}}.
 
This is a regression from the original port. Spark's \{{VariantUtil}}, which 
this class is derived from, passes the charset in both places:
 
{code:java}
return new String(value, start, length, StandardCharsets.UTF_8);
return new String(metadata, stringStart + offset, nextOffset - offset, 
StandardCharsets.UTF_8);
{code}
 
h3. Fix
Pass \{{StandardCharsets.UTF_8}} in both calls. Add a round-trip test with a 
non-ASCII string value and a non-ASCII field name.


> VARIANT string and object key decoding ignores the UTF-8 charset
> ----------------------------------------------------------------
>
>                 Key: FLINK-40305
>                 URL: https://issues.apache.org/jira/browse/FLINK-40305
>             Project: Flink
>          Issue Type: Bug
>          Components: API / Core
>            Reporter: Ramin Gharib
>            Assignee: Ramin Gharib
>            Priority: Major
>
> {{BinaryVariantUtil}} decodes variant strings and object field names without 
> passing a charset, so both fall back to the JVM default:
>  
> {code:java}
> // getString
> return new String(value, start, length);
> // getMetadataKey
> return new String(metadata, stringStart + offset, nextOffset - offset);
> {code}
>  
> Both are written as UTF-8 by {{{}BinaryVariantInternalBuilder{}}}, which uses 
> {{str.getBytes(StandardCharsets.UTF_8)}} for values and 
> {{key.getBytes(StandardCharsets.UTF_8)}} for keys, so the write and read 
> paths disagree.
>  
> The default charset is only guaranteed to be UTF-8 from Java 18 onwards (JEP 
> 400). On Java 11 and 17 it is platform dependent, so a variant holding 
> non-ASCII text can decode incorrectly. {{new String(byte[], ...)}} also 
> substitutes {{U+FFFD}} for malformed input rather than failing.
>  
> {{getMetadataKey}} is the more damaging of the two, since it decodes field 
> names. A mangled name makes {{getField(name)}} silently return null and 
> corrupts {{getFieldNames()}} and {{{}toJson(){}}}.
>  
> This is a regression from the original port. Spark's {{{}VariantUtil [1]{}}}, 
> which this class is derived from, passes the charset in both places:
>  
> {code:java}
> return new String(value, start, length, StandardCharsets.UTF_8);
> return new String(metadata, stringStart + offset, nextOffset - offset, 
> StandardCharsets.UTF_8);
> {code}
> h3. Fix
> Pass {{StandardCharsets.UTF_8}} in both calls. Add a round-trip test with a 
> non-ASCII string value and a non-ASCII field name.
>  
>  [1] 
> https://github.com/apache/spark/blob/master/common/variant/src/main/java/org/apache/spark/types/variant/VariantUtil.java#L617



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to