Claus Ibsen created CAMEL-25132:
-----------------------------------
Summary: camel-core - Routing engine hot-path performance
improvements
Key: CAMEL-25132
URL: https://issues.apache.org/jira/browse/CAMEL-25132
Project: Camel
Issue Type: Improvement
Components: camel-core
Reporter: Claus Ibsen
Fix For: 4.24.0
Hot-path performance analysis of the routing engine (per-exchange cost), with
measured numbers and prototyped fixes. This is an umbrella ticket: individual
items can be split into sub-tasks when work starts. Numbers were measured
against main at 927b570d022d (4.23.0-SNAPSHOT); re-verify against current code
before implementing.
Scope: per-exchange cost of the routing engine on main (4.23.0-SNAPSHOT, HEAD
927b570d022d). Startup is out of scope, and so is deprecated exchange pooling.
Method:
* A local harness measures ns/op and bytes/op on the calling thread
(ThreadMXBean), single-threaded through a ProducerTemplate.
* {{MtBench}} measures multi-threaded throughput.
* JFR allocation and CPU sampling show where time and memory go.
* Candidate fixes were prototyped as shadow classes placed ahead of the jars on
the classpath. The repository and {{\~/.m2}} were not changed.
* Concurrency numbers were rerun on Linux (Docker, aarch64, 8 vCPU, JDK 21). On
macOS, {{System.nanoTime()}} does not scale across threads because HotSpot uses
a global CAS to keep it monotonic, so macOS multi-thread numbers are misleading.
h2. Baseline (single thread, no JMX, via ProducerTemplate)
||scenario||ns/op||bytes/op||
|direct → 1 processor|400|1768|
|direct → 10 processors|1130|2264–2344|
|per extra route node|\~80 ns|\~55 B|
|per extra node with JMX (default stats)|\~136 ns|\~103 B|
|split of 100 items|34 µs|104 KB (\~1 KB and 340 ns per item)|
|10 nodes with {{maximumRedeliveries(3)}}|1467–1550|*8712*|
h2. Tier 1: high impact, measured
h3. 1. JMX processor statistics cap multi-threaded throughput (camel-management)
Linux, 10-node route, ops/s:
||setup||1 thread||4 threads||8 threads||
|no camel-management|942k|3.15M|5.49M|
|JMX statisticsLevel=Off|811k|2.78M|4.78M|
|JMX RoutesOnly|746k|1.58M|2.11M|
|JMX Default|*471k*|*622k*|*604k*|
With the default statistics level, throughput does not grow with thread count,
and single-threaded throughput is halved. Bisection (Linux, 8 threads):
||variant||ops/s||
|current|604k|
|counters prototype: StatisticCounter on LongAdder; StatisticValue/Delta write
only on change; mean computed on read; lastCompletedExchangeId updated at most
once per ms|1.87M (3.1×)|
|counters turned into no-ops|3.73M|
|no StopWatch either|4.66M|
Causes:
* {{ManagedPerformanceCounter.completedExchange}} / {{processExchange}} run for
every node, every exchange, and for three counters (processor, route, context):
** about 5 {{AtomicLong.getAndAdd}} calls, including inflight ++ and -- and
total processing time;
** about 6 volatile stores to shared objects (last, delta ×2, mean,
lastCompletedTimestamp, lastCompletedExchangeId);
** the mean is recomputed on every exchange.
* {{lastExchangeCompletedExchangeId = exchange.getExchangeId()}} forces UUID
generation for every exchange. Exchange ids are otherwise lazy.
* {{DefaultInstrumentationProcessor}} allocates a {{StopWatch}} and a state
{{Object[1]}} per node, and calls {{nanoTime}} twice per node. This costs about
28% single-threaded on this route, even though the stored value is truncated to
milliseconds.
Ideas:
* (a) Quick wins, already prototyped:
** LongAdder counters;
** write-if-changed for value statistics;
** compute the mean on read;
** throttle the last-exchange id and timestamp to once per ms;
** derive inflight as started − completed − failed.
* (b) Proper fix: striped per-counter cells, each holding all fields of one
counter in a padded object picked by thread probe. Updates become plain stores
in a mostly thread-local cache line; JMX reads aggregate.
* (c) Avoid two {{nanoTime}} calls per node: reuse the previous node's end
timestamp carried on the exchange, or sample timing for every Nth exchange at
processor level.
* (d) Consider making {{RoutesOnly}} the default for Spring Boot and Main. It
is a doc/default discussion, and it matters because camel-management on the
classpath silently switches processor-level stats on.
h3. 2. Redelivery-enabled error handler copies the whole Exchange at every node
{{RedeliveryErrorHandler.RedeliveryTask.prepare}} →
{{defensiveCopyExchangeIfNeeded}} → {{ExchangeHelper.createCopy}} runs on the
success path for every node when {{maximumRedeliveries != 0}}, {{retryWhile}}
is set, or any {{onException}} sets redeliveries. That is a very common
production setup. The copy is only ever read as {{original.getIn()}} in
{{redeliver()}}.
* Prototype: store {{exchange.getIn().copy()}} instead.
* Measured on 10 nodes: *8712 → 2408 bytes/exchange (−72%) and 1550 → 1158 ns
(−25%)*.
* Risk: low. Exchange properties were never restored on redelivery. Deprecate
the protected {{defensiveCopyExchangeIfNeeded}}; nothing overrides it in the
repo.
* Follow-ups:
** make RedeliveryTask implement AsyncCallback, which removes a capturing
lambda per delivery;
** build the full redelivery state only on the first failure.
h3. 3. Every exchange allocates a 336-byte EnumMap for internal properties
{{AbstractExchange.internalProperties = new
EnumMap<>(ExchangePropertyKey.class)}} has 69 keys, so the backing array is
{{Object[69]}}. That is 336 of the 488 bytes of {{new DefaultExchange}}, and it
is cloned by every copy (split, multicast, wiretap, recipient list, redelivery
copy).
* Creating the map lazily saves nothing, because {{TO_ENDPOINT}} is set on
every send ({{SendProcessor}}, {{DefaultProducerCache}}).
* Hybrid prototype: a dedicated field for {{TO_ENDPOINT}} plus a lazily created
EnumMap. Result: *−312 B per exchange* on normal routes, CPU neutral.
* A fully compact bitmap/packed-array map also saves about 260 B per split or
multicast item. My naive version cost about 10% CPU in split, so it needs
JMH-driven design before adopting.
* {{MessageSupport.traits}} (EnumMap over 2 constants) is allocated for every
message, and each exchange has 1–3 messages. Creating it lazily saved about 75
B per message. It is set only for redelivery/data-type traits.
* Compatibility: {{protected final EnumMap internalProperties}} and the
protected {{AbstractExchange}} constructor are the only API surface. They are
only used in camel-support and by the deprecated pooled exchange.
h3. 4. {{getEndpoint(String)}} runs a stream pipeline per call (camel-util
{{URISupport.textBlockToSingleLine}})
Every {{context.getEndpoint(uri)}} (ProducerTemplate with String URIs,
{{hasEndpoint}}, dynamic EIPs) runs {{uri.lines().forEach(...)}} with an
AtomicBoolean, a StringJoiner and per-line concatenation, even for
{{direct:foo}}. That is about 12% of all allocation in a ProducerTemplate loop.
* For a URI without {{\n}} or {{\r}}, the result is exactly {{uri.trim()}}, so
a fast path is a one-line change with identical behavior.
* Prototype together with #3 on a 1-processor route via template: 1768 → 928 B.
h3. Combined prototype (#2, #3 map, #3 traits, #4), bytes per exchange
||scenario||before||after||
|noop1|1768|928 (−47%)|
|noop10|2344|1504 (−36%)|
|typical route (filter/choice/simple/direct)|4064|3152 (−22%)|
|split 100|103,944|77,224 (−26%)|
|multicast ×3|6704|4944 (−26%)|
|redeliver10|8712|2408 (−72%)|
h3. 5. {{InputStream}}/{{Reader}} → String conversion: 29 KB of buffers per call
{{IOConverter.toString(InputStream, Exchange)}} goes through a BufferedReader,
an InputStreamReader and a StringBuilder.
* Verified: 1 KB body → *947 ns and 29,160 B*, versus 50 ns and 2,104 B for
{{new String(in.readAllBytes(), charset)}}.
* {{InputStreamCache}} (the default in-memory stream cache) takes the same
path, so every String view of an HTTP, JMS or file stream body pays this.
* Risk: low.
h3. 6. String → enum conversion scans the whole converter registry on every call
{{EnumTypeConverter}} → {{tcr.lookup}} → {{TypeResolverHelper.doLookup}}, which
never caches positive or negative results. It walks the class hierarchy twice
and scans all converters.
* Verified: {{convertTo(LoggingLevel, "INFO")}} costs *895 ns and 1,480 B*, and
gets worse with more converters on the classpath.
* It is hit by producers that read operation headers as enums and by predicates
on enum headers.
* Fix:
** cache lookups, including negative results, and clear the cache in
{{clearMisses()}};
** use a {{ClassValue}} for the enum-constant maps;
** avoid throwing on the {{tryConvertTo}} path.
h2. Tier 2: medium impact (code-verified, mostly not benchmarked)
* *Split/multicast sequential mode locking.* Each item takes about 5
lock/unlock pairs plus about 5 atomics:
** the AsyncCompletionService lock and {{signalAll}};
** the task {{tryLock}};
** two poll locks;
** the processor-wide {{doAggregateSync}} lock.
AQS release accounts for about 21–25% of CPU samples in single-threaded
split. Options: a lock-free sequential fast path, or skipping the processor
lock for built-in or per-exchange strategies (UseLatest, UseOriginal,
Grouped\*; keep it for user strategies).
* *Split per-item allocations:*
** a full {{DefaultUnitOfWork}} per sub-exchange, with an eager
{{ReentrantLock}} (about 48 B, CAS on every push/pop/getRoute);
** {{pairs.stream().filter(nonNull).toList()}} copy
({{MulticastProcessor:367}});
** {{PriorityQueue(capacity)}} sized to N even in sequential mode ({{:475}});
** a no-op release loop in {{doDone}} ({{:984}});
** a per-exchange {{UseOriginalAggregationStrategy}} plus a
{{ConcurrentHashMap}}.
* *Non-streaming split* builds every sub-exchange up front and keeps them all
alive until completion ({{Splitter:471}}, {{MulticastTask.pairs}}). Create
copies lazily over an eagerly read {{List<Object>}}. This changes when
{{onPrepare}} runs.
* *Throttler (TotalRequests, the default mode, and ConcurrentRequests)*
schedules and cancels a cleanup {{ScheduledFuture}} per exchange: about 3
executor-queue locks plus the per-key lock, and the rate expression is
evaluated under that lock. Reschedule only when the pending cleanup is near
firing. {{ThrottlePermit.compareTo}} calls {{currentTimeMillis}} twice per
comparison.
* *{{DefaultAsyncProcessorAwaitManager.process}}* allocates a
{{CountDownLatch}} and a lambda on every synchronous-over-asynchronous call
(ProducerTemplate, {{AsyncProcessorSupport.process(Exchange)}}, used by about
78 component call sites), even when the call completes synchronously. Use a
done flag with a lazily created latch.
* *{{DefaultProducerCache}} miss path:* {{ServicePool.acquire}} →
{{putIfAbsent}} → Caffeine or SimpleLRU {{compute}}, which locks and allocates
even when the entry exists. Try {{get}} first. {{lastUsedProducer}} also
ping-pongs between threads that use different endpoints.
* *Dynamic URIs re-normalized per exchange:* toD / enrich / wireTap
({{SendDynamicProcessor:337}}) and recipientList / routingSlip / dynamicRouter
({{ProcessorHelper:88}}) run {{normalizeUri}} on every exchange (49 ns / 216 B
for {{mock:x}}, about 1.3 KB with query params). Enrich always goes through
SendDynamicProcessor, even for constant URIs. A per-processor String →
NormalizedUri last-hit/LRU cache fixes this.
* *Enrich copies the message three times on InOut exchanges* with the default
strategy ({{Enricher:238/254/307}}). Skip the final {{copyResults}} when
{{aggregatedExchange == exchange}}.
* *Seda producer* takes the queue-reference lock twice per message
({{getCount}}, {{hasConsumers}}), shared by all producer threads. With
{{multipleConsumers}}, the consumer takes the endpoint lock per exchange.
* *{{PatternHelper.matchPattern}}* compiles a regex for every non-matching key.
{{removeHeaders("Camel*")}}, which the security checklist recommends, costs
about 930 ns and 7.6 KB per exchange with 12 headers.
{{AbstractExchange.removeProperties}} loops all 69 keys. Compile once per
pattern.
* *Simple language:*
** OGNL ({{${body.name}}}) builds a {{BeanExpression}} and copies the exchange
on every evaluation (361 ns, 3.8 KB);
** {{regex}}, {{in}}, {{is}} and {{range}} operators rebuild their right-hand
side every time;
** {{${date:...:pattern}}} creates a {{SimpleDateFormat}} per call (532 ns, 2
KB).
* *Type converter:*
** {{tryConvertTo}} misses are never cached;
** {{getBody(type)}} with a null body retries converting the Message itself,
walking its hierarchy (106 ns, 720 B);
** successful conversions allocate a {{TypeConvertible}} key and do the same
lookup twice.
* *{{CaseInsensitiveMap}}:*
** grows without bound under repeated put/remove of the same key (holes never
compacted);
** the copy constructor re-hashes every key (135 ns vs 34 ns for HashMap at 10
entries), which hits the first header write after every split, multicast or
wiretap copy;
** an empty map is 344 B.
* *{{ExchangeVariableRepository}}* is a full lifecycle service with a lock plus
a {{ConcurrentHashMap}}, created on the first {{setVariable}} and on every copy
of an exchange that has variables.
h2. Tier 3: engine core, small but broad
* *Per-node plumbing without JMX:* about 55 B per node (CamelInternalProcessor
{{AsyncAfterTask}} 24 B plus error handler {{SimpleTask}} about 32 B) and three
reactive-queue hops per node (SimpleTask, AsyncAfterTask, PipelineTask).
** Prototype: running the SimpleTask directly instead of {{scheduleMain}} gave
about 5–8% on 10 nodes. That touches reactive ordering, so it needs careful
review and tests.
* *{{DefaultReactiveExecutor.Worker}}:* {{queue}}, {{back}} and {{running}} are
{{volatile}}, although the worker is thread-confined through a ThreadLocal.
Removing volatile gave about 2–4%, which is within noise but consistent.
* *{{NodeHistoryAdvice}}* is always installed, with three setter calls before
and three after each node. It could be a single reference to a precomputed
immutable (id, label, source) holder.
* *{{CamelInternalProcessor}}* calls the deprecated
{{notifyExchangeAsyncProcessingStartedEvent}} after every node, because the
error handler always returns {{false}}, without the cached
{{isEventNotificationApplicable()}} guard. {{SendProcessor:270}} has the same
issue.
* *{{StopWatch.taken()}}* uses {{Duration.ofNanos(delta).toMillis()}}. Use
{{TimeUnit}}/division to avoid depending on escape analysis.
* *{{DefaultUnitOfWork}}:* an eager {{ReentrantLock}} per UoW, taken for
{{pushRoute}}, {{popRoute}}, {{getRoute}} and {{routeStackLevel}}. That is
about 4 or more CAS per route pass. The lock exists because CAMEL-20199
replaced {{synchronized}}. Consider lazy synchronizations or a cheaper guard.
* *{{RouteInflightRepositoryAdvice}}* does a {{routeCount.get(routeId)}} CHM
lookup twice per route pass. The LongAdder could be resolved once per route
start.
* *{{TracingAdvice}} / {{BacklogTracer}} standby (JBang, dev mode)* register an
event notifier that turns on ExchangeSending event creation for every send,
even while tracing is disabled. Worth checking the cost in dev mode.
h2. Failure path (only when failures are frequent)
* {{ExceptionPolicy.createRedeliveryPolicy}} clones and re-parses about 24
options with placeholder resolution on every failed attempt.
* {{DefaultExceptionPolicyStrategy}} allocates a TreeMap, two LinkedHashSets
and an ArrayList per exception, and re-splits route-scoped and context-scoped
policies each time.
* The "Failed delivery …" log strings are built before deciding whether to log.
A DLC also regex-sanitizes {{deadLetterUri}} on every call.
* {{RedeliveryPolicy}} uses a global lock and a shared {{Random}} for collision
avoidance; {{delayPattern}} is re-parsed on every redelivery.
h2. Possible correctness bugs found along the way (need tests)
* {{SendProcessor}} / {{DefaultProducerCache}} emit {{ExchangeSent}} only if
some notifier accepted {{ExchangeSending}}. {{SqlTraceDevConsole}}'s notifier
ignores Sending but wants Sent (CAMEL-14354 gating).
* A nested {{Splitter}} shares or removes the failure tracker through the
sub-exchange ({{Splitter.afterSend}}). The outer {{SplitResult}} counts can be
skewed.
* On the VIRTUAL thread-pool profile, the scheduled pool has no
{{removeOnCancelPolicy}}. Cancelled multicast/split timeouts keep the whole
{{MulticastTask}} alive until the delay expires.
h2. Suggested order
# Quick and low risk: #4 (text-block fast path), #2 (copy only the IN message),
#5 (IOConverter), the MessageSupport traits change, and {{StopWatch.taken}}.
# JMX counters: quick wins from #1(a), then the striped design.
# Internal-properties hybrid (#3), then JMH work on the compact map.
# Split/multicast sequential fast path and per-item allocations.
# Throttler cleanup scheduling, the await-manager latch, and the ProducerCache
miss path.
# Type-converter lookup caching (enum, try-miss) and Simple operator
precompilation.
# Reactive hops and the per-node plumbing (needs design review and benchmarks).
----
_Claude Code on behalf of davsclaus_
--
This message was sent by Atlassian Jira
(v8.20.10#820010)