Claus Ibsen created CAMEL-25078:
-----------------------------------
Summary: camel-xml-jaxp - XML type converters: fix bugs found in a
deep review
Key: CAMEL-25078
URL: https://issues.apache.org/jira/browse/CAMEL-25078
Project: Camel
Issue Type: Bug
Components: camel-core
Reporter: Claus Ibsen
A deep review of the XML type converters (camel-xml-jaxp) found the bugs below.
Each one was reproduced against 4.23.0-SNAPSHOT and has a test in
XmlConvertersEdgeCasesTest that fails without the fix. Several only fail with
the JDK StAX implementation, as the camel-core tests have Woodstox on the
classpath; the tests use the JDK implementation.
# *An XMLStreamReader converted to InputStream or Reader is cut at 16 KB by
readAllBytes().* A read of 0 bytes returned -1 (end of stream) instead of 0,
and readNBytes/readAllBytes make such a read.
# *An XMLStreamReader converted to InputStream with another charset than UTF-8
is empty or cut off (JDK StAX)*, as the writer was never flushed. This is used
with the charset of the exchange.
# *An XMLStreamReader converted to Reader (or String) fails on any attribute
without a namespace (JDK StAX).* The null guards added to the InputStream
variant in CAMEL-10120/CAMEL-12758 were missing in the Reader variant.
# *An XMLStreamReader positioned at an element (such as by nextTag) loses that
element* when converted to InputStream or Reader.
# *A text node in mixed content fails to convert to String with
ClassCastException* (such as the result of xpath /a/text() on
<a>foo<b/>bar</a>): the sibling after the text was cast to Text. A list of text
nodes also repeated the text of the siblings.
# *An attribute node converts to an empty String* instead of its value.
# *A file converted to XMLStreamReader or XMLEventReader ignores the encoding
of the XML declaration*, as the file variants used the default charset, while
the stream variants use the declaration (CAMEL-6779). Such as a split with stax
of a file in ISO-8859-1.
# *XmlLineNumberParser with root names fails when there are elements after the
root*, as it added a second document element (such as a bean after
camelContext); text before the root was added to it. The same copy in
camel-route-parser is fixed too.
*Not changed (for a later look)*
* A file converted to XMLStreamReader/XMLEventReader (or StAXSource from a
path) keeps the file open, as closing the reader does not close the stream.
* BytesSource.getReader() decodes with the platform default charset, which
parsers prefer over the bytes.
* The validator with failOnNullBody=false validates an empty document when the
body is not xml, which passes.
* XmlConverter.toStreamSource(String) and a few others use the platform default
charset instead of the declared encoding.
_Claude Code on behalf of Claus Ibsen_
--
This message was sent by Atlassian Jira
(v8.20.10#820010)