[ 
https://issues.apache.org/jira/browse/CAMEL-25038?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Claus Ibsen updated CAMEL-25038:
--------------------------------
    Fix Version/s: 4.23.0

> camel-disruptor - request/reply: a late reply is written into the caller's 
> exchange after ExchangeTimedOutException, and an interrupted wait reports 
> success (same bugs as CAMEL-24947 / CAMEL-24948 for seda)
> --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
>
>                 Key: CAMEL-25038
>                 URL: https://issues.apache.org/jira/browse/CAMEL-25038
>             Project: Camel
>          Issue Type: Bug
>          Components: camel-disruptor
>            Reporter: shashank
>            Priority: Minor
>             Fix For: 4.23.0
>
>
> {{DisruptorProducer.process()}} 
> ({{components/camel-disruptor/.../DisruptorProducer.java:73-149}}) waits for 
> a reply the same way {{SedaProducer}} does, and has the same two bugs that 
> CAMEL-24947 and CAMEL-24948 fix for camel-seda. The seda PRs #26793 (open) 
> and #26794 (merged) only change {{SedaProducer}}.
> # Timeout race. The consumer side ({{newOnCompletion}}, {{:152-172}}) does 
> {{if (latch.getCount() == 0) ignore; else 
> ExchangeHelper.copyResults(exchange, response); latch.countDown()}}. The 
> timeout branch ({{:108-123}}) does 
> {{exchange.setProperty(DISRUPTOR_IGNORE_EXCHANGE)}}, 
> {{exchange.setException(new ExchangeTimedOutException(...))}}, 
> {{latch.countDown()}} and then {{callback.done(true)}}. The check and the 
> copy are not atomic with the timeout branch. When the timeout fires while the 
> consumer is inside {{copyResults}}, the caller gets the reply body together 
> with {{ExchangeTimedOutException}}. The consumer then keeps writing into the 
> caller's exchange after the producer has returned: {{copyResults}} ends with 
> {{setException(source.getException())}}, which clears the timeout exception 
> while the caller's error handler or doCatch works on it. (Setting 
> {{DISRUPTOR_IGNORE_EXCHANGE}} on the caller's exchange has no effect on the 
> copy that is already in the ring.)
> # Interrupt. With {{timeout <= 0}} ({{:124-135}}), an interrupt of 
> {{latch.await()}} is logged at INFO and otherwise swallowed. The producer 
> then calls {{callback.done(true)}} with no exception, so the caller sees a 
> successful exchange whose "reply" is its own request. The real reply is 
> written into the exchange later. With {{timeout > 0}} ({{:101-106}}), an 
> interrupt leaves {{done}} false and is reported as a timeout, which is 
> acceptable, but that path still has race 1.
> *Reproduction* (4.23.0-SNAPSHOT). As in the seda ticket, the pause is a 
> {{SafeCopyProperty}} set by the consumer route; {{copyResults}} calls it 
> after the message copy and before {{setException}}:
> {noformat}
> paused:    template.send(disruptor:b?timeout=300, InOut), consumer reply copy 
> paused inside onDone
>              returned to caller after 337 ms: body=reply, 
> exception=ExchangeTimedOutException ...
>              same exchange 400 ms later:        body=reply, exception=null
> natural:   400 requests, timeout=20 ms, consumer takes 18-22 ms, no pause:
>              timeouts=238, returned with the reply body AND 
> ExchangeTimedOutException=25
> interrupt: InOut timeout=0, caller thread interrupted while waiting:
>              returned body=request, exception=null, failed=false
>              600 ms later: body=reply, exception=null
> {noformat}
> TLA+: the seda producer/consumer model, adapted to {{DisruptorProducer}}: 
> {{pw_timeout_race}} violates {{NoWriteAfterReturn}}, {{pw_timeout_view}} 
> violates {{CallerViewStable}}, and {{pw_notimeout_interrupt}} violates 
> {{ConsistentOutcome}}. The fixed variant {{pw_fixed}} satisfies all of them 
> plus {{ProducerReturns}}.
> *Proposed fix:* the same as for seda (CAMEL-24947 / CAMEL-24948):
> * Let exactly one side own the result through an {{AtomicBoolean}} shared by 
> the producer and the onDone synchronization. onDone claims it before copying 
> and ignores the reply if the claim fails. On timeout, the producer claims it. 
> If the producer wins, it sets {{ExchangeTimedOutException}}. If it loses, the 
> consumer is copying the reply, so it waits for the latch without a timeout 
> and returns the reply.
> * On {{InterruptedException}}, restore the interrupt flag, claim the result, 
> and fail the exchange (for example with a {{RejectedExecutionException}} or 
> {{CamelExchangeException}} "interrupted while waiting for reply") instead of 
> reporting success.
> _Filed with Claude Code on behalf of allthingssecurity._



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to