shashank created CAMEL-25126:
--------------------------------

             Summary: camel-syslog - the syslog data format decodes every byte 
as ISO-8859-1, so non-ASCII text (UTF-8 MSG and structured data values) is 
corrupted, and the endpoint encoding option does not help
                 Key: CAMEL-25126
                 URL: https://issues.apache.org/jira/browse/CAMEL-25126
             Project: Camel
          Issue Type: Bug
          Components: camel-syslog
            Reporter: shashank


{{SyslogConverter.parseMessage}} turns each byte of the message into one 
character with {{(char) (byteBuffer.get() & 0xff)}}, which is an ISO-8859-1 
decoding. {{SyslogDataFormat.unmarshal}} first reads the input stream into a 
String with the exchange charset (UTF-8 by default), then calls 
{{body.getBytes()}} (the JVM default charset, UTF-8 since Java 18) and gives 
those bytes to {{parseMessage}}. So on a current JVM every character outside 
US-ASCII comes back as two to four wrong characters, with no error:

{noformat}
sent (UTF-8):  café Grüße
received:      café GrüÃ?e        (? is the control character U+009F; é 
becomes U+00C3 U+00A9, ü U+00C3 U+00BC, ß U+00C3 U+009F)
{noformat}

RFC 5424 says that MSG SHOULD be Unicode encoded as UTF-8 (section 6.4), that a 
MSG which starts with the UTF-8 BOM is UTF-8 (MSG-UTF8 = BOM UTF-8-STRING), and 
that a PARAM-VALUE of the structured data is a UTF-8 string (section 6.3.3). So 
an accent, an umlaut or a non-Latin script in the log text or in a structured 
data value is corrupted in the Camel message. The BOM itself becomes the three 
characters {{}} at the start of the log message. RFC 3164 messages go 
through the same code. The header fields HOSTNAME, APP-NAME, PROCID and MSGID 
are US-ASCII by definition and are not the problem.

This is the normal consumer path, not only the type converter. With the route 
of the documentation, 
{{from("netty:udp://0.0.0.0:10514?sync=false&allowDefaultCodec=false").unmarshal().syslog()}},
 a UTF-8 datagram {{<14>Oct  1 10:00:00 host café Grüße}} gives the log message 
{{café Grü�e}} (? being U+009F). Setting the charset on the endpoint does 
not help: with {{encoding=ISO-8859-1}} on the netty endpoint and a Latin-1 
datagram, the result is the same wrong text, because the data format re-encodes 
the String with the JVM default charset before it parses it.

The data format is not symmetric either: {{marshal}} writes the text with 
{{getBytes()}}, and {{unmarshal}} of those bytes returns a different text, so 
{{unmarshal(marshal(m)).getLogMessage()}} is not {{m.getLogMessage()}} for any 
non-ASCII text, for both {{SyslogMessage}} and {{Rfc5424SyslogMessage}}.

h3. Reproduction

{code:java}
SyslogMessage m = SyslogConverter.toSyslogMessage("<14>Oct  1 10:00:00 host 
café");
// m.getLogMessage() is "café", expected "café"
{code}

Also reproduced end to end with a netty UDP route and {{unmarshal().syslog()}}, 
with and without the {{encoding}} option, on main (three runs). A small formal 
model (Lean 4) of the decoding shows that the result is wrong for every text 
with at least one character outside US-ASCII, that US-ASCII text is unchanged, 
and that decoding the bytes of each field with UTF-8 gives every Unicode text 
back.

h3. Affected versions

The byte-to-char loops are the same in 2.14.0, 2.16.0, 3.0.0, 4.0.0, 4.14.0, 
4.18.0, 4.22.0 and main. The corruption on the consumer path depends on the JVM 
default charset: with a default of ISO-8859-1 (possible on Java 17) Latin-1 
characters come out right; with UTF-8, the default since Java 18, they do not.

h3. Proposed fix

* {{parseMessage}} decodes each field with a charset, instead of one character 
per byte. A new {{parseMessage(byte\[\], Charset)}}; {{parseMessage(byte\[\])}} 
uses UTF-8.
* A MSG that starts with the UTF-8 BOM is decoded as UTF-8, without the BOM, 
whatever the charset.
* {{SyslogDataFormat.unmarshal}} parses the bytes as received and decodes them 
with the charset of the exchange ({{CamelCharsetName}}, which the {{encoding}} 
option of camel-netty and camel-mina sets), UTF-8 by default. {{marshal}} 
writes with the same charset.
* {{toSyslogMessage(String)}} uses UTF-8 for its internal bytes, so a String is 
parsed without loss.

US-ASCII messages are parsed exactly as before. Wherever the current code gives 
the right text (a JVM whose default charset is ISO-8859-1 and characters in 
that range), the fix gives the same text. The change is visible to users who 
worked around the wrong text, so it needs a note in the upgrade guide.

Duplicate check (2026-09-29): JIRA "syslog" with "UTF-8", "charset", 
"encoding", "unicode" or "BOM", and "SyslogConverter": only CAMEL-8687 
(structured data with spaces) and CAMEL-14369 (message without PRI); nothing 
about the character decoding. No open pull request touches camel-syslog.

_Filed with Claude Code on behalf of allthingssecurity._




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to