shashank created CAMEL-25129:
--------------------------------
Summary: camel-resilience4j - Circuit Breaker EIP: with timeout
enabled, the bulkhead permit is released when the call times out while the call
keeps running, so bulkheadMaxConcurrentCalls is not enforced
Key: CAMEL-25129
URL: https://issues.apache.org/jira/browse/CAMEL-25129
Project: Camel
Issue Type: Bug
Components: camel-core
Reporter: shashank
With {{timeoutEnabled}} the protected processor runs on the timeout executor
and the caller waits for it with the resilience4j {{TimeLimiter}}. The bulkhead
is applied around that wait:
* synchronous: {{TimeLimiter.decorateFutureSupplier}} over a supplier of
{{CompletableFuture.supplyAsync(task, executorService)}}, then
{{Bulkhead.decorateCallable(bulkhead, callable)}}
({{ResilienceProcessor.processSync}});
* asynchronous ({{asynchronous=true}}):
{{TimeLimiter.decorateCompletionStage(...)}}, then
{{Bulkhead.decorateCompletionStage(...)}} ({{processAsync}}).
So the bulkhead permit is held by the caller and released when the
{{TimeLimiter}} gives up. The work itself is not stopped (a
{{CompletableFuture}} is not interrupted by {{cancel}}, and the task keeps
running on the timeout executor), so after a timeout the call still runs but no
longer holds a permit. Every further message gets a permit, times out and
leaves another call running. The number of calls that run at the same time
against the protected service is bounded by the timeout thread pool (default
profile: 10 to 20 threads, 1000 queued), not by {{bulkheadMaxConcurrentCalls}}.
That is the case the bulkhead exists for: a slow downstream. With a bulkhead of
1 and a timeout, a slow service gets one new concurrent call per timeout period
and message instead of one call at a time. Without timeout the bulkhead works,
because the call runs on the caller thread inside the permit.
The order comes from CAMEL-17095 ("Using timeout and bulkhead together does not
work"), a one-line fix so that the bulkhead decorates the time limited callable
instead of the bare task; the order itself was not discussed. Resilience4j
documents the order of its aspects as
Retry(CircuitBreaker(RateLimiter(TimeLimiter(Bulkhead(function))))), and its
README example for asynchronous calls applies the bulkhead first, then the time
limiter, then the circuit breaker: the bulkhead inside the time limiter, so
that the work holds the permit until it ends.
camel-microprofile-fault-tolerance already behaves that way: SmallRye applies
the bulkhead inside the timeout, and its synchronous timeout returns only when
the invocation has returned, so there the permit is held for the whole call.
h3. Reproduction
A route with
{{circuitBreaker().resilience4jConfiguration().bulkheadEnabled(true).bulkheadMaxConcurrentCalls(1).timeoutEnabled(true).timeoutDuration(200)}},
a protected processor that blocks on a latch (a slow downstream) and counts
the calls that run at the same time, and a fallback. Three messages are sent
one after another; each returns with the fallback after 200 ms.
* synchronous: 3 protected calls are running at the same time after the third
message (maximum 3, fallback 3), 3 of 3 runs.
* {{asynchronous=true}}: the same, 3 of 3 runs.
* control, no timeout, three concurrent senders: at most 1 call runs, the other
two are rejected by the bulkhead and get the fallback, 3 of 3 runs.
A TLA+ model with callers, the time limiter, the worker and the bulkhead checks
that no more than {{maxConcurrentCalls}} protected calls execute at the same
time. It is violated in 6 steps (call 1 starts, times out and releases its
permit while running, call 2 gets a permit and runs). It holds on the control
without timeout, and holds with the bulkhead inside the time limiter, for 3
calls and 1 or 2 permits, with every call completing.
h3. Proposed fix
Take the bulkhead permit around the work, inside the time limiter:
* synchronous: wrap the supplier of {{CompletableFuture.supplyAsync(task,
executorService)}} with {{Bulkhead.decorateCompletionStage(bulkhead, ...)}},
and apply {{TimeLimiter.decorateFutureSupplier}} over it; the outer
{{Bulkhead.decorateCallable}} is only used without timeout;
* asynchronous: apply {{Bulkhead.decorateCompletionStage}} before
{{TimeLimiter.decorateCompletionStage}}.
The permit is then released when the work completes. The caller still waits for
a permit up to {{bulkheadMaxWaitDuration}} before the timeout starts. A full
bulkhead still reaches the fallback as a {{BulkheadFullException}} (through the
future), so the fallback, the {{CamelCircuitBreakerResponseRejected}} property
and the bulkhead-rejected counter stay the same.
Behaviour change for the upgrade guide: while calls that timed out are still
running, further calls are rejected by the bulkhead (fallback,
{{CamelCircuitBreakerResponseRejected}} true) where they used to be started; a
call that never ends keeps its permit.
camel-microprofile-fault-tolerance is not affected (see above) and is not
changed.
Tests: the reproduction above with a latch, synchronous and asynchronous,
checking that the second and third calls are rejected by the bulkhead and that
the permit is released when the slow call ends.
Affected: all versions since CAMEL-17095 (3.11.4 and 3.13.0) for the
synchronous path, and the asynchronous mode since it was added in 4.22.0
(CAMEL-24209).
Duplicate check (2026-09-29): JIRA text "bulkhead" (9 issues: CAMEL-17095,
CAMEL-16173 bulkhead not applied at all, CAMEL-24819 metrics, CAMEL-24209, the
review umbrellas CAMEL-24136 and CAMEL-24137, and older fault-tolerance ones),
"circuit breaker" with "fallback" since July 2026 (9 issues): none about the
permit being released on timeout. GitHub pull requests "resilience4j bulkhead",
"circuit breaker bulkhead timeout", "exchangeWriteGuard": dependency bumps,
#24977 (CAMEL-24209) and #24797 (CAMEL-24134), not this.
_Filed with Claude Code on behalf of allthingssecurity._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)