Claus Ibsen created CAMEL-25132:
-----------------------------------

             Summary: camel-core - Routing engine hot-path performance 
improvements
                 Key: CAMEL-25132
                 URL: https://issues.apache.org/jira/browse/CAMEL-25132
             Project: Camel
          Issue Type: Improvement
          Components: camel-core
            Reporter: Claus Ibsen
             Fix For: 4.24.0


Hot-path performance analysis of the routing engine (per-exchange cost), with 
measured numbers and prototyped fixes. This is an umbrella ticket: individual 
items can be split into sub-tasks when work starts. Numbers were measured 
against main at 927b570d022d (4.23.0-SNAPSHOT); re-verify against current code 
before implementing.

Scope: per-exchange cost of the routing engine on main (4.23.0-SNAPSHOT, HEAD 
927b570d022d). Startup is out of scope, and so is deprecated exchange pooling.

Method:
* A local harness measures ns/op and bytes/op on the calling thread 
(ThreadMXBean), single-threaded through a ProducerTemplate.
* {{MtBench}} measures multi-threaded throughput.
* JFR allocation and CPU sampling show where time and memory go.
* Candidate fixes were prototyped as shadow classes placed ahead of the jars on 
the classpath. The repository and {{\~/.m2}} were not changed.
* Concurrency numbers were rerun on Linux (Docker, aarch64, 8 vCPU, JDK 21). On 
macOS, {{System.nanoTime()}} does not scale across threads because HotSpot uses 
a global CAS to keep it monotonic, so macOS multi-thread numbers are misleading.

h2. Baseline (single thread, no JMX, via ProducerTemplate)

||scenario||ns/op||bytes/op||
|direct → 1 processor|400|1768|
|direct → 10 processors|1130|2264–2344|
|per extra route node|\~80 ns|\~55 B|
|per extra node with JMX (default stats)|\~136 ns|\~103 B|
|split of 100 items|34 µs|104 KB (\~1 KB and 340 ns per item)|
|10 nodes with {{maximumRedeliveries(3)}}|1467–1550|*8712*|

h2. Tier 1: high impact, measured

h3. 1. JMX processor statistics cap multi-threaded throughput (camel-management)
Linux, 10-node route, ops/s:

||setup||1 thread||4 threads||8 threads||
|no camel-management|942k|3.15M|5.49M|
|JMX statisticsLevel=Off|811k|2.78M|4.78M|
|JMX RoutesOnly|746k|1.58M|2.11M|
|JMX Default|*471k*|*622k*|*604k*|

With the default statistics level, throughput does not grow with thread count, 
and single-threaded throughput is halved. Bisection (Linux, 8 threads):

||variant||ops/s||
|current|604k|
|counters prototype: StatisticCounter on LongAdder; StatisticValue/Delta write 
only on change; mean computed on read; lastCompletedExchangeId updated at most 
once per ms|1.87M (3.1×)|
|counters turned into no-ops|3.73M|
|no StopWatch either|4.66M|

Causes:
* {{ManagedPerformanceCounter.completedExchange}} / {{processExchange}} run for 
every node, every exchange, and for three counters (processor, route, context):
** about 5 {{AtomicLong.getAndAdd}} calls, including inflight ++ and -- and 
total processing time;
** about 6 volatile stores to shared objects (last, delta ×2, mean, 
lastCompletedTimestamp, lastCompletedExchangeId);
** the mean is recomputed on every exchange.
* {{lastExchangeCompletedExchangeId = exchange.getExchangeId()}} forces UUID 
generation for every exchange. Exchange ids are otherwise lazy.
* {{DefaultInstrumentationProcessor}} allocates a {{StopWatch}} and a state 
{{Object[1]}} per node, and calls {{nanoTime}} twice per node. This costs about 
28% single-threaded on this route, even though the stored value is truncated to 
milliseconds.

Ideas:
* (a) Quick wins, already prototyped:
** LongAdder counters;
** write-if-changed for value statistics;
** compute the mean on read;
** throttle the last-exchange id and timestamp to once per ms;
** derive inflight as started − completed − failed.
* (b) Proper fix: striped per-counter cells, each holding all fields of one 
counter in a padded object picked by thread probe. Updates become plain stores 
in a mostly thread-local cache line; JMX reads aggregate.
* (c) Avoid two {{nanoTime}} calls per node: reuse the previous node's end 
timestamp carried on the exchange, or sample timing for every Nth exchange at 
processor level.
* (d) Consider making {{RoutesOnly}} the default for Spring Boot and Main. It 
is a doc/default discussion, and it matters because camel-management on the 
classpath silently switches processor-level stats on.

h3. 2. Redelivery-enabled error handler copies the whole Exchange at every node
{{RedeliveryErrorHandler.RedeliveryTask.prepare}} → 
{{defensiveCopyExchangeIfNeeded}} → {{ExchangeHelper.createCopy}} runs on the 
success path for every node when {{maximumRedeliveries != 0}}, {{retryWhile}} 
is set, or any {{onException}} sets redeliveries. That is a very common 
production setup. The copy is only ever read as {{original.getIn()}} in 
{{redeliver()}}.
* Prototype: store {{exchange.getIn().copy()}} instead.
* Measured on 10 nodes: *8712 → 2408 bytes/exchange (−72%) and 1550 → 1158 ns 
(−25%)*.
* Risk: low. Exchange properties were never restored on redelivery. Deprecate 
the protected {{defensiveCopyExchangeIfNeeded}}; nothing overrides it in the 
repo.
* Follow-ups:
** make RedeliveryTask implement AsyncCallback, which removes a capturing 
lambda per delivery;
** build the full redelivery state only on the first failure.

h3. 3. Every exchange allocates a 336-byte EnumMap for internal properties
{{AbstractExchange.internalProperties = new 
EnumMap<>(ExchangePropertyKey.class)}} has 69 keys, so the backing array is 
{{Object[69]}}. That is 336 of the 488 bytes of {{new DefaultExchange}}, and it 
is cloned by every copy (split, multicast, wiretap, recipient list, redelivery 
copy).
* Creating the map lazily saves nothing, because {{TO_ENDPOINT}} is set on 
every send ({{SendProcessor}}, {{DefaultProducerCache}}).
* Hybrid prototype: a dedicated field for {{TO_ENDPOINT}} plus a lazily created 
EnumMap. Result: *−312 B per exchange* on normal routes, CPU neutral.
* A fully compact bitmap/packed-array map also saves about 260 B per split or 
multicast item. My naive version cost about 10% CPU in split, so it needs 
JMH-driven design before adopting.
* {{MessageSupport.traits}} (EnumMap over 2 constants) is allocated for every 
message, and each exchange has 1–3 messages. Creating it lazily saved about 75 
B per message. It is set only for redelivery/data-type traits.
* Compatibility: {{protected final EnumMap internalProperties}} and the 
protected {{AbstractExchange}} constructor are the only API surface. They are 
only used in camel-support and by the deprecated pooled exchange.

h3. 4. {{getEndpoint(String)}} runs a stream pipeline per call (camel-util 
{{URISupport.textBlockToSingleLine}})
Every {{context.getEndpoint(uri)}} (ProducerTemplate with String URIs, 
{{hasEndpoint}}, dynamic EIPs) runs {{uri.lines().forEach(...)}} with an 
AtomicBoolean, a StringJoiner and per-line concatenation, even for 
{{direct:foo}}. That is about 12% of all allocation in a ProducerTemplate loop.
* For a URI without {{\n}} or {{\r}}, the result is exactly {{uri.trim()}}, so 
a fast path is a one-line change with identical behavior.
* Prototype together with #3 on a 1-processor route via template: 1768 → 928 B.

h3. Combined prototype (#2, #3 map, #3 traits, #4), bytes per exchange

||scenario||before||after||
|noop1|1768|928 (−47%)|
|noop10|2344|1504 (−36%)|
|typical route (filter/choice/simple/direct)|4064|3152 (−22%)|
|split 100|103,944|77,224 (−26%)|
|multicast ×3|6704|4944 (−26%)|
|redeliver10|8712|2408 (−72%)|

h3. 5. {{InputStream}}/{{Reader}} → String conversion: 29 KB of buffers per call
{{IOConverter.toString(InputStream, Exchange)}} goes through a BufferedReader, 
an InputStreamReader and a StringBuilder.
* Verified: 1 KB body → *947 ns and 29,160 B*, versus 50 ns and 2,104 B for 
{{new String(in.readAllBytes(), charset)}}.
* {{InputStreamCache}} (the default in-memory stream cache) takes the same 
path, so every String view of an HTTP, JMS or file stream body pays this.
* Risk: low.

h3. 6. String → enum conversion scans the whole converter registry on every call
{{EnumTypeConverter}} → {{tcr.lookup}} → {{TypeResolverHelper.doLookup}}, which 
never caches positive or negative results. It walks the class hierarchy twice 
and scans all converters.
* Verified: {{convertTo(LoggingLevel, "INFO")}} costs *895 ns and 1,480 B*, and 
gets worse with more converters on the classpath.
* It is hit by producers that read operation headers as enums and by predicates 
on enum headers.
* Fix:
** cache lookups, including negative results, and clear the cache in 
{{clearMisses()}};
** use a {{ClassValue}} for the enum-constant maps;
** avoid throwing on the {{tryConvertTo}} path.

h2. Tier 2: medium impact (code-verified, mostly not benchmarked)

* *Split/multicast sequential mode locking.* Each item takes about 5 
lock/unlock pairs plus about 5 atomics:
** the AsyncCompletionService lock and {{signalAll}};
** the task {{tryLock}};
** two poll locks;
** the processor-wide {{doAggregateSync}} lock.

  AQS release accounts for about 21–25% of CPU samples in single-threaded 
split. Options: a lock-free sequential fast path, or skipping the processor 
lock for built-in or per-exchange strategies (UseLatest, UseOriginal, 
Grouped\*; keep it for user strategies).
* *Split per-item allocations:*
** a full {{DefaultUnitOfWork}} per sub-exchange, with an eager 
{{ReentrantLock}} (about 48 B, CAS on every push/pop/getRoute);
** {{pairs.stream().filter(nonNull).toList()}} copy 
({{MulticastProcessor:367}});
** {{PriorityQueue(capacity)}} sized to N even in sequential mode ({{:475}});
** a no-op release loop in {{doDone}} ({{:984}});
** a per-exchange {{UseOriginalAggregationStrategy}} plus a 
{{ConcurrentHashMap}}.
* *Non-streaming split* builds every sub-exchange up front and keeps them all 
alive until completion ({{Splitter:471}}, {{MulticastTask.pairs}}). Create 
copies lazily over an eagerly read {{List<Object>}}. This changes when 
{{onPrepare}} runs.
* *Throttler (TotalRequests, the default mode, and ConcurrentRequests)* 
schedules and cancels a cleanup {{ScheduledFuture}} per exchange: about 3 
executor-queue locks plus the per-key lock, and the rate expression is 
evaluated under that lock. Reschedule only when the pending cleanup is near 
firing. {{ThrottlePermit.compareTo}} calls {{currentTimeMillis}} twice per 
comparison.
* *{{DefaultAsyncProcessorAwaitManager.process}}* allocates a 
{{CountDownLatch}} and a lambda on every synchronous-over-asynchronous call 
(ProducerTemplate, {{AsyncProcessorSupport.process(Exchange)}}, used by about 
78 component call sites), even when the call completes synchronously. Use a 
done flag with a lazily created latch.
* *{{DefaultProducerCache}} miss path:* {{ServicePool.acquire}} → 
{{putIfAbsent}} → Caffeine or SimpleLRU {{compute}}, which locks and allocates 
even when the entry exists. Try {{get}} first. {{lastUsedProducer}} also 
ping-pongs between threads that use different endpoints.
* *Dynamic URIs re-normalized per exchange:* toD / enrich / wireTap 
({{SendDynamicProcessor:337}}) and recipientList / routingSlip / dynamicRouter 
({{ProcessorHelper:88}}) run {{normalizeUri}} on every exchange (49 ns / 216 B 
for {{mock:x}}, about 1.3 KB with query params). Enrich always goes through 
SendDynamicProcessor, even for constant URIs. A per-processor String → 
NormalizedUri last-hit/LRU cache fixes this.
* *Enrich copies the message three times on InOut exchanges* with the default 
strategy ({{Enricher:238/254/307}}). Skip the final {{copyResults}} when 
{{aggregatedExchange == exchange}}.
* *Seda producer* takes the queue-reference lock twice per message 
({{getCount}}, {{hasConsumers}}), shared by all producer threads. With 
{{multipleConsumers}}, the consumer takes the endpoint lock per exchange.
* *{{PatternHelper.matchPattern}}* compiles a regex for every non-matching key. 
{{removeHeaders("Camel*")}}, which the security checklist recommends, costs 
about 930 ns and 7.6 KB per exchange with 12 headers. 
{{AbstractExchange.removeProperties}} loops all 69 keys. Compile once per 
pattern.
* *Simple language:*
** OGNL ({{${body.name}}}) builds a {{BeanExpression}} and copies the exchange 
on every evaluation (361 ns, 3.8 KB);
** {{regex}}, {{in}}, {{is}} and {{range}} operators rebuild their right-hand 
side every time;
** {{${date:...:pattern}}} creates a {{SimpleDateFormat}} per call (532 ns, 2 
KB).
* *Type converter:*
** {{tryConvertTo}} misses are never cached;
** {{getBody(type)}} with a null body retries converting the Message itself, 
walking its hierarchy (106 ns, 720 B);
** successful conversions allocate a {{TypeConvertible}} key and do the same 
lookup twice.
* *{{CaseInsensitiveMap}}:*
** grows without bound under repeated put/remove of the same key (holes never 
compacted);
** the copy constructor re-hashes every key (135 ns vs 34 ns for HashMap at 10 
entries), which hits the first header write after every split, multicast or 
wiretap copy;
** an empty map is 344 B.
* *{{ExchangeVariableRepository}}* is a full lifecycle service with a lock plus 
a {{ConcurrentHashMap}}, created on the first {{setVariable}} and on every copy 
of an exchange that has variables.

h2. Tier 3: engine core, small but broad

* *Per-node plumbing without JMX:* about 55 B per node (CamelInternalProcessor 
{{AsyncAfterTask}} 24 B plus error handler {{SimpleTask}} about 32 B) and three 
reactive-queue hops per node (SimpleTask, AsyncAfterTask, PipelineTask).
** Prototype: running the SimpleTask directly instead of {{scheduleMain}} gave 
about 5–8% on 10 nodes. That touches reactive ordering, so it needs careful 
review and tests.
* *{{DefaultReactiveExecutor.Worker}}:* {{queue}}, {{back}} and {{running}} are 
{{volatile}}, although the worker is thread-confined through a ThreadLocal. 
Removing volatile gave about 2–4%, which is within noise but consistent.
* *{{NodeHistoryAdvice}}* is always installed, with three setter calls before 
and three after each node. It could be a single reference to a precomputed 
immutable (id, label, source) holder.
* *{{CamelInternalProcessor}}* calls the deprecated 
{{notifyExchangeAsyncProcessingStartedEvent}} after every node, because the 
error handler always returns {{false}}, without the cached 
{{isEventNotificationApplicable()}} guard. {{SendProcessor:270}} has the same 
issue.
* *{{StopWatch.taken()}}* uses {{Duration.ofNanos(delta).toMillis()}}. Use 
{{TimeUnit}}/division to avoid depending on escape analysis.
* *{{DefaultUnitOfWork}}:* an eager {{ReentrantLock}} per UoW, taken for 
{{pushRoute}}, {{popRoute}}, {{getRoute}} and {{routeStackLevel}}. That is 
about 4 or more CAS per route pass. The lock exists because CAMEL-20199 
replaced {{synchronized}}. Consider lazy synchronizations or a cheaper guard.
* *{{RouteInflightRepositoryAdvice}}* does a {{routeCount.get(routeId)}} CHM 
lookup twice per route pass. The LongAdder could be resolved once per route 
start.
* *{{TracingAdvice}} / {{BacklogTracer}} standby (JBang, dev mode)* register an 
event notifier that turns on ExchangeSending event creation for every send, 
even while tracing is disabled. Worth checking the cost in dev mode.

h2. Failure path (only when failures are frequent)

* {{ExceptionPolicy.createRedeliveryPolicy}} clones and re-parses about 24 
options with placeholder resolution on every failed attempt.
* {{DefaultExceptionPolicyStrategy}} allocates a TreeMap, two LinkedHashSets 
and an ArrayList per exception, and re-splits route-scoped and context-scoped 
policies each time.
* The "Failed delivery …" log strings are built before deciding whether to log. 
A DLC also regex-sanitizes {{deadLetterUri}} on every call.
* {{RedeliveryPolicy}} uses a global lock and a shared {{Random}} for collision 
avoidance; {{delayPattern}} is re-parsed on every redelivery.

h2. Possible correctness bugs found along the way (need tests)

* {{SendProcessor}} / {{DefaultProducerCache}} emit {{ExchangeSent}} only if 
some notifier accepted {{ExchangeSending}}. {{SqlTraceDevConsole}}'s notifier 
ignores Sending but wants Sent (CAMEL-14354 gating).
* A nested {{Splitter}} shares or removes the failure tracker through the 
sub-exchange ({{Splitter.afterSend}}). The outer {{SplitResult}} counts can be 
skewed.
* On the VIRTUAL thread-pool profile, the scheduled pool has no 
{{removeOnCancelPolicy}}. Cancelled multicast/split timeouts keep the whole 
{{MulticastTask}} alive until the delay expires.

h2. Suggested order

# Quick and low risk: #4 (text-block fast path), #2 (copy only the IN message), 
#5 (IOConverter), the MessageSupport traits change, and {{StopWatch.taken}}.
# JMX counters: quick wins from #1(a), then the striped design.
# Internal-properties hybrid (#3), then JMH work on the compact map.
# Split/multicast sequential fast path and per-item allocations.
# Throttler cleanup scheduling, the await-manager latch, and the ProducerCache 
miss path.
# Type-converter lookup caching (enum, try-miss) and Simple operator 
precompilation.
# Reactive hops and the per-node plumbing (needs design review and benchmarks).

----
_Claude Code on behalf of davsclaus_



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to