[
https://issues.apache.org/jira/browse/CAMEL-25038?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Claus Ibsen updated CAMEL-25038:
--------------------------------
Fix Version/s: 4.23.0
> camel-disruptor - request/reply: a late reply is written into the caller's
> exchange after ExchangeTimedOutException, and an interrupted wait reports
> success (same bugs as CAMEL-24947 / CAMEL-24948 for seda)
> --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
>
> Key: CAMEL-25038
> URL: https://issues.apache.org/jira/browse/CAMEL-25038
> Project: Camel
> Issue Type: Bug
> Components: camel-disruptor
> Reporter: shashank
> Priority: Minor
> Fix For: 4.23.0
>
>
> {{DisruptorProducer.process()}}
> ({{components/camel-disruptor/.../DisruptorProducer.java:73-149}}) waits for
> a reply the same way {{SedaProducer}} does, and has the same two bugs that
> CAMEL-24947 and CAMEL-24948 fix for camel-seda. The seda PRs #26793 (open)
> and #26794 (merged) only change {{SedaProducer}}.
> # Timeout race. The consumer side ({{newOnCompletion}}, {{:152-172}}) does
> {{if (latch.getCount() == 0) ignore; else
> ExchangeHelper.copyResults(exchange, response); latch.countDown()}}. The
> timeout branch ({{:108-123}}) does
> {{exchange.setProperty(DISRUPTOR_IGNORE_EXCHANGE)}},
> {{exchange.setException(new ExchangeTimedOutException(...))}},
> {{latch.countDown()}} and then {{callback.done(true)}}. The check and the
> copy are not atomic with the timeout branch. When the timeout fires while the
> consumer is inside {{copyResults}}, the caller gets the reply body together
> with {{ExchangeTimedOutException}}. The consumer then keeps writing into the
> caller's exchange after the producer has returned: {{copyResults}} ends with
> {{setException(source.getException())}}, which clears the timeout exception
> while the caller's error handler or doCatch works on it. (Setting
> {{DISRUPTOR_IGNORE_EXCHANGE}} on the caller's exchange has no effect on the
> copy that is already in the ring.)
> # Interrupt. With {{timeout <= 0}} ({{:124-135}}), an interrupt of
> {{latch.await()}} is logged at INFO and otherwise swallowed. The producer
> then calls {{callback.done(true)}} with no exception, so the caller sees a
> successful exchange whose "reply" is its own request. The real reply is
> written into the exchange later. With {{timeout > 0}} ({{:101-106}}), an
> interrupt leaves {{done}} false and is reported as a timeout, which is
> acceptable, but that path still has race 1.
> *Reproduction* (4.23.0-SNAPSHOT). As in the seda ticket, the pause is a
> {{SafeCopyProperty}} set by the consumer route; {{copyResults}} calls it
> after the message copy and before {{setException}}:
> {noformat}
> paused: template.send(disruptor:b?timeout=300, InOut), consumer reply copy
> paused inside onDone
> returned to caller after 337 ms: body=reply,
> exception=ExchangeTimedOutException ...
> same exchange 400 ms later: body=reply, exception=null
> natural: 400 requests, timeout=20 ms, consumer takes 18-22 ms, no pause:
> timeouts=238, returned with the reply body AND
> ExchangeTimedOutException=25
> interrupt: InOut timeout=0, caller thread interrupted while waiting:
> returned body=request, exception=null, failed=false
> 600 ms later: body=reply, exception=null
> {noformat}
> TLA+: the seda producer/consumer model, adapted to {{DisruptorProducer}}:
> {{pw_timeout_race}} violates {{NoWriteAfterReturn}}, {{pw_timeout_view}}
> violates {{CallerViewStable}}, and {{pw_notimeout_interrupt}} violates
> {{ConsistentOutcome}}. The fixed variant {{pw_fixed}} satisfies all of them
> plus {{ProducerReturns}}.
> *Proposed fix:* the same as for seda (CAMEL-24947 / CAMEL-24948):
> * Let exactly one side own the result through an {{AtomicBoolean}} shared by
> the producer and the onDone synchronization. onDone claims it before copying
> and ignores the reply if the claim fails. On timeout, the producer claims it.
> If the producer wins, it sets {{ExchangeTimedOutException}}. If it loses, the
> consumer is copying the reply, so it waits for the latch without a timeout
> and returns the reply.
> * On {{InterruptedException}}, restore the interrupt flag, claim the result,
> and fail the exchange (for example with a {{RejectedExecutionException}} or
> {{CamelExchangeException}} "interrupted while waiting for reply") instead of
> reporting success.
> _Filed with Claude Code on behalf of allthingssecurity._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)