shashank created CAMEL-25152:
--------------------------------
Summary: camel-bindy - fixed-length marshal pads and clips fields
in UTF-16 chars while unmarshal counts code points (or graphemes), so a record
with an emoji or other non-BMP character is written one pad short and read back
shifted
Key: CAMEL-25152
URL: https://issues.apache.org/jira/browse/CAMEL-25152
Project: Camel
Issue Type: Bug
Components: camel-bindy
Reporter: shashank
Since CAMEL-14521 (3.1.0) the fixed-length unmarshal cuts a record with
{{UnicodeHelper}}: a field of {{length = 5}} is 5 code points, or 5 graphemes
with {{@FixedLengthRecord(countGrapheme = true)}}. Marshal was not changed:
{{BindyFixedLengthFactory.generateFixedLengthPositionMap}} pads a field with
{{length - result.length()}} padding characters and clips it with
{{result.substring(0, length)}}, both in UTF-16 chars. The record length check
of {{BindyFixedLengthDataFormat.createModel}} also uses {{String.length()}}.
For text in the Basic Multilingual Plane with {{countGrapheme = false}} the two
counts are the same. A character outside the BMP (emoji, CJK Extension B, many
historic scripts) is one code point but two UTF-16 chars, and with
{{countGrapheme = true}} a base letter with a combining mark ({{e}} + U+0301)
is one grapheme but two chars. Such a field is written one padding character
short per extra char, and unmarshal of the record then takes the missing
characters from the next field, so every following field is shifted:
{noformat}
model: @FixedLengthRecord(length = 12), fields of length 5, 4, 3 (align R, trim
= true), values "ok😀", "x", "y"
marshal today: " ok😀 x y" (4 code points for the first field, 11 in all)
unmarshal: "ok😀 ", "x ", "y"
fixed marshal: " ok😀 x y" -> "ok😀", "x", "y"
the correctly padded record " ok😀 x y" (12 code points) written by another
system:
unmarshal today: IllegalArgumentException: Size of the record: 13 is not equal
to the value provided in the model: 12
{noformat}
With {{clip = true}} the clip can also cut a surrogate pair in half and write
an invalid character. With {{countGrapheme = true}} (added for exactly this
kind of text) a decomposed {{café}} is read back as {{café }} and the next
fields are shifted the same way.
h3. Reproduction
A unit test ({{BindyFixedLengthMarshalUnicodeTest}}) marshals and unmarshals
records with {{ok😀}}, two CJK Extension B characters and a decomposed {{café}}
(with {{countGrapheme = true}}), and clips {{abcd😀f}} to 5; all four cases fail
on main (three runs) and pass with the fix. A small formal model (Lean 4) shows
that padding by code points gives every record back whose values fit their
lengths, that any value with a non-BMP character followed by another field is
read back with extra characters today, and that for BMP text the fix writes
exactly what is written today.
h3. Affected versions
Unmarshal has counted code points since 3.1.0 (CAMEL-14521) and marshal still
counts UTF-16 chars in 3.1.0, 4.0.0, 4.22.0 and main. Before 3.1.0 both counted
UTF-16 chars.
h3. Proposed fix
* In {{generateFixedLengthPositionMap}} measure and clip the formatted value
with {{UnicodeHelper}} and the same method as unmarshal ({{CODEPOINTS}}, or
{{GRAPHEME}} with {{countGrapheme}}): pad with {{length -
unicodeResult.length()}} and clip with {{unicodeResult.substring(0, length)}}.
* In {{createModel}} compare and cut the record length
({{ignoreTrailingChars}}) with the same count.
Text in the BMP without combining marks is written and read exactly as today.
Duplicate check (2026-09-30): JIRA "bindy" with "grapheme", "unicode",
"surrogate", "emoji", "codepoint": CAMEL-14521 and CAMEL-14825 (unmarshal only,
fixed), CAMEL-14085 (bytes vs chars, Won't Fix); nothing about marshal. No open
pull request touches the fixed-length marshal.
_Filed with Claude Code on behalf of allthingssecurity._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)