Claus Ibsen created CAMEL-25078:
-----------------------------------

             Summary: camel-xml-jaxp - XML type converters: fix bugs found in a 
deep review
                 Key: CAMEL-25078
                 URL: https://issues.apache.org/jira/browse/CAMEL-25078
             Project: Camel
          Issue Type: Bug
          Components: camel-core
            Reporter: Claus Ibsen


A deep review of the XML type converters (camel-xml-jaxp) found the bugs below. 
Each one was reproduced against 4.23.0-SNAPSHOT and has a test in 
XmlConvertersEdgeCasesTest that fails without the fix. Several only fail with 
the JDK StAX implementation, as the camel-core tests have Woodstox on the 
classpath; the tests use the JDK implementation.

# *An XMLStreamReader converted to InputStream or Reader is cut at 16 KB by 
readAllBytes().* A read of 0 bytes returned -1 (end of stream) instead of 0, 
and readNBytes/readAllBytes make such a read.
# *An XMLStreamReader converted to InputStream with another charset than UTF-8 
is empty or cut off (JDK StAX)*, as the writer was never flushed. This is used 
with the charset of the exchange.
# *An XMLStreamReader converted to Reader (or String) fails on any attribute 
without a namespace (JDK StAX).* The null guards added to the InputStream 
variant in CAMEL-10120/CAMEL-12758 were missing in the Reader variant.
# *An XMLStreamReader positioned at an element (such as by nextTag) loses that 
element* when converted to InputStream or Reader.
# *A text node in mixed content fails to convert to String with 
ClassCastException* (such as the result of xpath /a/text() on 
<a>foo<b/>bar</a>): the sibling after the text was cast to Text. A list of text 
nodes also repeated the text of the siblings.
# *An attribute node converts to an empty String* instead of its value.
# *A file converted to XMLStreamReader or XMLEventReader ignores the encoding 
of the XML declaration*, as the file variants used the default charset, while 
the stream variants use the declaration (CAMEL-6779). Such as a split with stax 
of a file in ISO-8859-1.
# *XmlLineNumberParser with root names fails when there are elements after the 
root*, as it added a second document element (such as a bean after 
camelContext); text before the root was added to it. The same copy in 
camel-route-parser is fixed too.

*Not changed (for a later look)*
* A file converted to XMLStreamReader/XMLEventReader (or StAXSource from a 
path) keeps the file open, as closing the reader does not close the stream.
* BytesSource.getReader() decodes with the platform default charset, which 
parsers prefer over the bytes.
* The validator with failOnNullBody=false validates an empty document when the 
body is not xml, which passes.
* XmlConverter.toStreamSource(String) and a few others use the platform default 
charset instead of the declared encoding.

_Claude Code on behalf of Claus Ibsen_




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to