shashank created CAMEL-25217:
--------------------------------

             Summary: camel-platform-http-vertx - a String response body is 
always written as UTF-8, even when the response Content-Type declares another 
charset (for example charset=ISO-8859-1), so the client decodes it wrongly
                 Key: CAMEL-25217
                 URL: https://issues.apache.org/jira/browse/CAMEL-25217
             Project: Camel
          Issue Type: Bug
          Components: camel-platform-http-vertx
            Reporter: shashank


{{VertxPlatformHttpSupport.toHttpResponse}} copies the {{Content-Type}} of the 
message, with its {{charset}} parameter, to the response. {{writeResponse}} 
then writes a {{String}} body with

{code:java}
} else if (body instanceof String string) {
    ctx.end(string);
{code}

and Vert.x encodes a {{String}} as UTF-8, whatever the declared charset. A 
client that decodes the response with the charset of the {{Content-Type}} (as 
it should, RFC 9110 8.3.1) gets mojibake for every non-ASCII character. Two 
common cases:

* a route that sets {{Content-Type: text/plain; charset=ISO-8859-1}} and a 
{{String}} body ({{setBody(constant(...))}}, {{simple}}, {{transform}}): 
{{Grüße aus Köln}} (14 characters) is sent as 17 bytes of UTF-8 labelled 
ISO-8859-1, and read as {{Grü...}} (every non-ASCII character becomes two);
* an echo, or any route that keeps the request {{Content-Type}}: a request with 
{{charset=ISO-8859-1}} is read correctly (the consumer sets 
{{CamelCharsetName}} from it), but the reply, labelled with the same 
{{Content-Type}}, is UTF-8.

{{byte[]}}, {{InputStream}} and {{Buffer}} bodies are not affected, and neither 
is ASCII text, which is why the existing tests pass. The type converter 
{{VertxBufferConverter.toBuffer(String, Exchange)}} already uses the charset of 
the message {{Content-Type}} first, then {{CamelCharsetName}} (CAMEL-16282); 
the {{String}} branch of {{writeResponse}} bypasses it. (camel-servlet encodes 
a {{String}} in the exchange charset, {{CamelCharsetName}} or UTF-8, and only 
in non-chunked mode also sets the response character encoding to match; so it 
gets the echo case right but not a route that only sets the {{Content-Type}}.)

h3. Reproduction

{code:java}
from("platform-http:/latin1")
    .setHeader(Exchange.CONTENT_TYPE, constant("text/plain; 
charset=ISO-8859-1"))
    .setBody(constant("Grüße aus Köln"));

from("platform-http:/echo")
    .convertBodyTo(String.class);
{code}

{{GET /latin1}} returns {{Content-Type: text/plain; charset=ISO-8859-1}} with 
the UTF-8 bytes of the text (17 bytes instead of 14). {{POST /echo}} with 
{{Content-Type: text/plain; charset=ISO-8859-1}} and the 14 ISO-8859-1 bytes 
returns 17 bytes. A unit test with both routes fails on main; the control with 
{{Content-Type: text/plain}} (no charset, UTF-8 expected) passes. A small 
formal model (Lean 4) shows that for every text with a character outside ASCII 
what a client reads with the declared ISO-8859-1 charset is longer than the 
text, and that writing in the declared charset returns every text the charset 
can represent.

h3. Proposed fix

Write the {{String}} in the charset of the response {{Content-Type}} when it 
declares one (and the JVM supports it), otherwise keep UTF-8 as today:

{code:java}
} else if (body instanceof String string) {
    // write the text in the charset declared by the response Content-Type 
(vert.x writes a String as UTF-8)
    String charset = responseCharset(ctx);
    if (charset != null) {
        ctx.end(Buffer.buffer(string, charset));
    } else {
        ctx.end(string);
    }
    promise.complete();

private static String responseCharset(RoutingContext ctx) {
    String contentType = ctx.response().headers().get("Content-Type");
    if (contentType == null) {
        return null;
    }
    String charset = IOHelper.getCharsetNameFromContentType(contentType);
    if (charset == null || charset.isEmpty()) {
        return null;
    }
    try {
        return Charset.isSupported(charset) ? charset : null;
    } catch (IllegalCharsetNameException e) {
        return null;
    }
}
{code}

Only responses that declare a charset other than UTF-8 change, and for them the 
bytes now match the header. A response without a charset parameter stays UTF-8 
(using {{CamelCharsetName}} there as well, like camel-servlet, would change the 
bytes of responses whose header does not say so, so it is left out). With the 
fix the new test and the whole camel-platform-http-vertx suite pass (150 tests, 
1 skipped, camel-core built from the same commit). Because the bytes of these 
responses change, the change gets an upgrade guide note (a client that ignored 
the declared charset and read UTF-8 must now use the declared one; characters 
the declared charset cannot represent are written as {{?}}).

Duplicate check (2026-09-30): JIRA "platform-http" with "charset", "encoding" 
or "UTF-8", and "VertxPlatformHttpSupport": CAMEL-16282, CAMEL-16756, 
CAMEL-22125, CAMEL-23320 (Spring Boot starter, binary data), none about a 
String body. No pull request about it.

_Filed with Claude Code on behalf of allthingssecurity._




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to