shashank created CAMEL-25152:
--------------------------------

             Summary: camel-bindy - fixed-length marshal pads and clips fields 
in UTF-16 chars while unmarshal counts code points (or graphemes), so a record 
with an emoji or other non-BMP character is written one pad short and read back 
shifted
                 Key: CAMEL-25152
                 URL: https://issues.apache.org/jira/browse/CAMEL-25152
             Project: Camel
          Issue Type: Bug
          Components: camel-bindy
            Reporter: shashank


Since CAMEL-14521 (3.1.0) the fixed-length unmarshal cuts a record with 
{{UnicodeHelper}}: a field of {{length = 5}} is 5 code points, or 5 graphemes 
with {{@FixedLengthRecord(countGrapheme = true)}}. Marshal was not changed: 
{{BindyFixedLengthFactory.generateFixedLengthPositionMap}} pads a field with 
{{length - result.length()}} padding characters and clips it with 
{{result.substring(0, length)}}, both in UTF-16 chars. The record length check 
of {{BindyFixedLengthDataFormat.createModel}} also uses {{String.length()}}.

For text in the Basic Multilingual Plane with {{countGrapheme = false}} the two 
counts are the same. A character outside the BMP (emoji, CJK Extension B, many 
historic scripts) is one code point but two UTF-16 chars, and with 
{{countGrapheme = true}} a base letter with a combining mark ({{e}} + U+0301) 
is one grapheme but two chars. Such a field is written one padding character 
short per extra char, and unmarshal of the record then takes the missing 
characters from the next field, so every following field is shifted:

{noformat}
model: @FixedLengthRecord(length = 12), fields of length 5, 4, 3 (align R, trim 
= true), values "ok😀", "x", "y"
marshal today: " ok😀   x  y"   (4 code points for the first field, 11 in all)
unmarshal:     "ok😀 ", "x ", "y"
fixed marshal: "  ok😀   x  y"   -> "ok😀", "x", "y"

the correctly padded record "  ok😀   x  y" (12 code points) written by another 
system:
unmarshal today: IllegalArgumentException: Size of the record: 13 is not equal 
to the value provided in the model: 12
{noformat}

With {{clip = true}} the clip can also cut a surrogate pair in half and write 
an invalid character. With {{countGrapheme = true}} (added for exactly this 
kind of text) a decomposed {{café}} is read back as {{café }} and the next 
fields are shifted the same way.

h3. Reproduction

A unit test ({{BindyFixedLengthMarshalUnicodeTest}}) marshals and unmarshals 
records with {{ok😀}}, two CJK Extension B characters and a decomposed {{café}} 
(with {{countGrapheme = true}}), and clips {{abcd😀f}} to 5; all four cases fail 
on main (three runs) and pass with the fix. A small formal model (Lean 4) shows 
that padding by code points gives every record back whose values fit their 
lengths, that any value with a non-BMP character followed by another field is 
read back with extra characters today, and that for BMP text the fix writes 
exactly what is written today.

h3. Affected versions

Unmarshal has counted code points since 3.1.0 (CAMEL-14521) and marshal still 
counts UTF-16 chars in 3.1.0, 4.0.0, 4.22.0 and main. Before 3.1.0 both counted 
UTF-16 chars.

h3. Proposed fix

* In {{generateFixedLengthPositionMap}} measure and clip the formatted value 
with {{UnicodeHelper}} and the same method as unmarshal ({{CODEPOINTS}}, or 
{{GRAPHEME}} with {{countGrapheme}}): pad with {{length - 
unicodeResult.length()}} and clip with {{unicodeResult.substring(0, length)}}.
* In {{createModel}} compare and cut the record length 
({{ignoreTrailingChars}}) with the same count.

Text in the BMP without combining marks is written and read exactly as today.

Duplicate check (2026-09-30): JIRA "bindy" with "grapheme", "unicode", 
"surrogate", "emoji", "codepoint": CAMEL-14521 and CAMEL-14825 (unmarshal only, 
fixed), CAMEL-14085 (bytes vs chars, Won't Fix); nothing about marshal. No open 
pull request touches the fixed-length marshal.

_Filed with Claude Code on behalf of allthingssecurity._




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to