[ 
https://issues.apache.org/jira/browse/CAMEL-25132?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18120810#comment-18120810
 ] 

Brijesh Thakkar commented on CAMEL-25132:
-----------------------------------------

[~davsclaus] I really want to work on this issue

Will raise PR for it

Thank youu

> camel-core - Routing engine hot-path performance improvements
> -------------------------------------------------------------
>
>                 Key: CAMEL-25132
>                 URL: https://issues.apache.org/jira/browse/CAMEL-25132
>             Project: Camel
>          Issue Type: Improvement
>          Components: camel-core
>            Reporter: Claus Ibsen
>            Priority: Major
>             Fix For: 4.24.0
>
>
> Hot-path performance analysis of the routing engine (per-exchange cost), with 
> measured numbers and prototyped fixes. This is an umbrella ticket: individual 
> items can be split into sub-tasks when work starts. Numbers were measured 
> against main at 927b570d022d (4.23.0-SNAPSHOT); re-verify against current 
> code before implementing.
> Scope: per-exchange cost of the routing engine on main (4.23.0-SNAPSHOT, HEAD 
> 927b570d022d). Startup is out of scope, and so is deprecated exchange pooling.
> Method:
> * A local harness measures ns/op and bytes/op on the calling thread 
> (ThreadMXBean), single-threaded through a ProducerTemplate.
> * {{MtBench}} measures multi-threaded throughput.
> * JFR allocation and CPU sampling show where time and memory go.
> * Candidate fixes were prototyped as shadow classes placed ahead of the jars 
> on the classpath. The repository and {{\~/.m2}} were not changed.
> * Concurrency numbers were rerun on Linux (Docker, aarch64, 8 vCPU, JDK 21). 
> On macOS, {{System.nanoTime()}} does not scale across threads because HotSpot 
> uses a global CAS to keep it monotonic, so macOS multi-thread numbers are 
> misleading.
> h2. Baseline (single thread, no JMX, via ProducerTemplate)
> ||scenario||ns/op||bytes/op||
> |direct → 1 processor|400|1768|
> |direct → 10 processors|1130|2264–2344|
> |per extra route node|\~80 ns|\~55 B|
> |per extra node with JMX (default stats)|\~136 ns|\~103 B|
> |split of 100 items|34 µs|104 KB (\~1 KB and 340 ns per item)|
> |10 nodes with {{maximumRedeliveries(3)}}|1467–1550|*8712*|
> h2. Tier 1: high impact, measured
> h3. 1. JMX processor statistics cap multi-threaded throughput 
> (camel-management)
> Linux, 10-node route, ops/s:
> ||setup||1 thread||4 threads||8 threads||
> |no camel-management|942k|3.15M|5.49M|
> |JMX statisticsLevel=Off|811k|2.78M|4.78M|
> |JMX RoutesOnly|746k|1.58M|2.11M|
> |JMX Default|*471k*|*622k*|*604k*|
> With the default statistics level, throughput does not grow with thread 
> count, and single-threaded throughput is halved. Bisection (Linux, 8 threads):
> ||variant||ops/s||
> |current|604k|
> |counters prototype: StatisticCounter on LongAdder; StatisticValue/Delta 
> write only on change; mean computed on read; lastCompletedExchangeId updated 
> at most once per ms|1.87M (3.1×)|
> |counters turned into no-ops|3.73M|
> |no StopWatch either|4.66M|
> Causes:
> * {{ManagedPerformanceCounter.completedExchange}} / {{processExchange}} run 
> for every node, every exchange, and for three counters (processor, route, 
> context):
> ** about 5 {{AtomicLong.getAndAdd}} calls, including inflight ++ and -- and 
> total processing time;
> ** about 6 volatile stores to shared objects (last, delta ×2, mean, 
> lastCompletedTimestamp, lastCompletedExchangeId);
> ** the mean is recomputed on every exchange.
> * {{lastExchangeCompletedExchangeId = exchange.getExchangeId()}} forces UUID 
> generation for every exchange. Exchange ids are otherwise lazy.
> * {{DefaultInstrumentationProcessor}} allocates a {{StopWatch}} and a state 
> {{Object[1]}} per node, and calls {{nanoTime}} twice per node. This costs 
> about 28% single-threaded on this route, even though the stored value is 
> truncated to milliseconds.
> Ideas:
> * (a) Quick wins, already prototyped:
> ** LongAdder counters;
> ** write-if-changed for value statistics;
> ** compute the mean on read;
> ** throttle the last-exchange id and timestamp to once per ms;
> ** derive inflight as started − completed − failed.
> * (b) Proper fix: striped per-counter cells, each holding all fields of one 
> counter in a padded object picked by thread probe. Updates become plain 
> stores in a mostly thread-local cache line; JMX reads aggregate.
> * (c) Avoid two {{nanoTime}} calls per node: reuse the previous node's end 
> timestamp carried on the exchange, or sample timing for every Nth exchange at 
> processor level.
> * (d) Consider making {{RoutesOnly}} the default for Spring Boot and Main. It 
> is a doc/default discussion, and it matters because camel-management on the 
> classpath silently switches processor-level stats on.
> h3. 2. Redelivery-enabled error handler copies the whole Exchange at every 
> node
> {{RedeliveryErrorHandler.RedeliveryTask.prepare}} → 
> {{defensiveCopyExchangeIfNeeded}} → {{ExchangeHelper.createCopy}} runs on the 
> success path for every node when {{maximumRedeliveries != 0}}, {{retryWhile}} 
> is set, or any {{onException}} sets redeliveries. That is a very common 
> production setup. The copy is only ever read as {{original.getIn()}} in 
> {{redeliver()}}.
> * Prototype: store {{exchange.getIn().copy()}} instead.
> * Measured on 10 nodes: *8712 → 2408 bytes/exchange (−72%) and 1550 → 1158 ns 
> (−25%)*.
> * Risk: low. Exchange properties were never restored on redelivery. Deprecate 
> the protected {{defensiveCopyExchangeIfNeeded}}; nothing overrides it in the 
> repo.
> * Follow-ups:
> ** make RedeliveryTask implement AsyncCallback, which removes a capturing 
> lambda per delivery;
> ** build the full redelivery state only on the first failure.
> h3. 3. Every exchange allocates a 336-byte EnumMap for internal properties
> {{AbstractExchange.internalProperties = new 
> EnumMap<>(ExchangePropertyKey.class)}} has 69 keys, so the backing array is 
> {{Object[69]}}. That is 336 of the 488 bytes of {{new DefaultExchange}}, and 
> it is cloned by every copy (split, multicast, wiretap, recipient list, 
> redelivery copy).
> * Creating the map lazily saves nothing, because {{TO_ENDPOINT}} is set on 
> every send ({{SendProcessor}}, {{DefaultProducerCache}}).
> * Hybrid prototype: a dedicated field for {{TO_ENDPOINT}} plus a lazily 
> created EnumMap. Result: *−312 B per exchange* on normal routes, CPU neutral.
> * A fully compact bitmap/packed-array map also saves about 260 B per split or 
> multicast item. My naive version cost about 10% CPU in split, so it needs 
> JMH-driven design before adopting.
> * {{MessageSupport.traits}} (EnumMap over 2 constants) is allocated for every 
> message, and each exchange has 1–3 messages. Creating it lazily saved about 
> 75 B per message. It is set only for redelivery/data-type traits.
> * Compatibility: {{protected final EnumMap internalProperties}} and the 
> protected {{AbstractExchange}} constructor are the only API surface. They are 
> only used in camel-support and by the deprecated pooled exchange.
> h3. 4. {{getEndpoint(String)}} runs a stream pipeline per call (camel-util 
> {{URISupport.textBlockToSingleLine}})
> Every {{context.getEndpoint(uri)}} (ProducerTemplate with String URIs, 
> {{hasEndpoint}}, dynamic EIPs) runs {{uri.lines().forEach(...)}} with an 
> AtomicBoolean, a StringJoiner and per-line concatenation, even for 
> {{direct:foo}}. That is about 12% of all allocation in a ProducerTemplate 
> loop.
> * For a URI without {{\n}} or {{\r}}, the result is exactly {{uri.trim()}}, 
> so a fast path is a one-line change with identical behavior.
> * Prototype together with #3 on a 1-processor route via template: 1768 → 928 
> B.
> h3. Combined prototype (#2, #3 map, #3 traits, #4), bytes per exchange
> ||scenario||before||after||
> |noop1|1768|928 (−47%)|
> |noop10|2344|1504 (−36%)|
> |typical route (filter/choice/simple/direct)|4064|3152 (−22%)|
> |split 100|103,944|77,224 (−26%)|
> |multicast ×3|6704|4944 (−26%)|
> |redeliver10|8712|2408 (−72%)|
> h3. 5. {{InputStream}}/{{Reader}} → String conversion: 29 KB of buffers per 
> call
> {{IOConverter.toString(InputStream, Exchange)}} goes through a 
> BufferedReader, an InputStreamReader and a StringBuilder.
> * Verified: 1 KB body → *947 ns and 29,160 B*, versus 50 ns and 2,104 B for 
> {{new String(in.readAllBytes(), charset)}}.
> * {{InputStreamCache}} (the default in-memory stream cache) takes the same 
> path, so every String view of an HTTP, JMS or file stream body pays this.
> * Risk: low.
> h3. 6. String → enum conversion scans the whole converter registry on every 
> call
> {{EnumTypeConverter}} → {{tcr.lookup}} → {{TypeResolverHelper.doLookup}}, 
> which never caches positive or negative results. It walks the class hierarchy 
> twice and scans all converters.
> * Verified: {{convertTo(LoggingLevel, "INFO")}} costs *895 ns and 1,480 B*, 
> and gets worse with more converters on the classpath.
> * It is hit by producers that read operation headers as enums and by 
> predicates on enum headers.
> * Fix:
> ** cache lookups, including negative results, and clear the cache in 
> {{clearMisses()}};
> ** use a {{ClassValue}} for the enum-constant maps;
> ** avoid throwing on the {{tryConvertTo}} path.
> h2. Tier 2: medium impact (code-verified, mostly not benchmarked)
> * *Split/multicast sequential mode locking.* Each item takes about 5 
> lock/unlock pairs plus about 5 atomics:
> ** the AsyncCompletionService lock and {{signalAll}};
> ** the task {{tryLock}};
> ** two poll locks;
> ** the processor-wide {{doAggregateSync}} lock.
>   AQS release accounts for about 21–25% of CPU samples in single-threaded 
> split. Options: a lock-free sequential fast path, or skipping the processor 
> lock for built-in or per-exchange strategies (UseLatest, UseOriginal, 
> Grouped\*; keep it for user strategies).
> * *Split per-item allocations:*
> ** a full {{DefaultUnitOfWork}} per sub-exchange, with an eager 
> {{ReentrantLock}} (about 48 B, CAS on every push/pop/getRoute);
> ** {{pairs.stream().filter(nonNull).toList()}} copy 
> ({{MulticastProcessor:367}});
> ** {{PriorityQueue(capacity)}} sized to N even in sequential mode ({{:475}});
> ** a no-op release loop in {{doDone}} ({{:984}});
> ** a per-exchange {{UseOriginalAggregationStrategy}} plus a 
> {{ConcurrentHashMap}}.
> * *Non-streaming split* builds every sub-exchange up front and keeps them all 
> alive until completion ({{Splitter:471}}, {{MulticastTask.pairs}}). Create 
> copies lazily over an eagerly read {{List<Object>}}. This changes when 
> {{onPrepare}} runs.
> * *Throttler (TotalRequests, the default mode, and ConcurrentRequests)* 
> schedules and cancels a cleanup {{ScheduledFuture}} per exchange: about 3 
> executor-queue locks plus the per-key lock, and the rate expression is 
> evaluated under that lock. Reschedule only when the pending cleanup is near 
> firing. {{ThrottlePermit.compareTo}} calls {{currentTimeMillis}} twice per 
> comparison.
> * *{{DefaultAsyncProcessorAwaitManager.process}}* allocates a 
> {{CountDownLatch}} and a lambda on every synchronous-over-asynchronous call 
> (ProducerTemplate, {{AsyncProcessorSupport.process(Exchange)}}, used by about 
> 78 component call sites), even when the call completes synchronously. Use a 
> done flag with a lazily created latch.
> * *{{DefaultProducerCache}} miss path:* {{ServicePool.acquire}} → 
> {{putIfAbsent}} → Caffeine or SimpleLRU {{compute}}, which locks and 
> allocates even when the entry exists. Try {{get}} first. {{lastUsedProducer}} 
> also ping-pongs between threads that use different endpoints.
> * *Dynamic URIs re-normalized per exchange:* toD / enrich / wireTap 
> ({{SendDynamicProcessor:337}}) and recipientList / routingSlip / 
> dynamicRouter ({{ProcessorHelper:88}}) run {{normalizeUri}} on every exchange 
> (49 ns / 216 B for {{mock:x}}, about 1.3 KB with query params). Enrich always 
> goes through SendDynamicProcessor, even for constant URIs. A per-processor 
> String → NormalizedUri last-hit/LRU cache fixes this.
> * *Enrich copies the message three times on InOut exchanges* with the default 
> strategy ({{Enricher:238/254/307}}). Skip the final {{copyResults}} when 
> {{aggregatedExchange == exchange}}.
> * *Seda producer* takes the queue-reference lock twice per message 
> ({{getCount}}, {{hasConsumers}}), shared by all producer threads. With 
> {{multipleConsumers}}, the consumer takes the endpoint lock per exchange.
> * *{{PatternHelper.matchPattern}}* compiles a regex for every non-matching 
> key. {{removeHeaders("Camel*")}}, which the security checklist recommends, 
> costs about 930 ns and 7.6 KB per exchange with 12 headers. 
> {{AbstractExchange.removeProperties}} loops all 69 keys. Compile once per 
> pattern.
> * *Simple language:*
> ** OGNL ({{${body.name}}}) builds a {{BeanExpression}} and copies the 
> exchange on every evaluation (361 ns, 3.8 KB);
> ** {{regex}}, {{in}}, {{is}} and {{range}} operators rebuild their right-hand 
> side every time;
> ** {{${date:...:pattern}}} creates a {{SimpleDateFormat}} per call (532 ns, 2 
> KB).
> * *Type converter:*
> ** {{tryConvertTo}} misses are never cached;
> ** {{getBody(type)}} with a null body retries converting the Message itself, 
> walking its hierarchy (106 ns, 720 B);
> ** successful conversions allocate a {{TypeConvertible}} key and do the same 
> lookup twice.
> * *{{CaseInsensitiveMap}}:*
> ** grows without bound under repeated put/remove of the same key (holes never 
> compacted);
> ** the copy constructor re-hashes every key (135 ns vs 34 ns for HashMap at 
> 10 entries), which hits the first header write after every split, multicast 
> or wiretap copy;
> ** an empty map is 344 B.
> * *{{ExchangeVariableRepository}}* is a full lifecycle service with a lock 
> plus a {{ConcurrentHashMap}}, created on the first {{setVariable}} and on 
> every copy of an exchange that has variables.
> h2. Tier 3: engine core, small but broad
> * *Per-node plumbing without JMX:* about 55 B per node 
> (CamelInternalProcessor {{AsyncAfterTask}} 24 B plus error handler 
> {{SimpleTask}} about 32 B) and three reactive-queue hops per node 
> (SimpleTask, AsyncAfterTask, PipelineTask).
> ** Prototype: running the SimpleTask directly instead of {{scheduleMain}} 
> gave about 5–8% on 10 nodes. That touches reactive ordering, so it needs 
> careful review and tests.
> * *{{DefaultReactiveExecutor.Worker}}:* {{queue}}, {{back}} and {{running}} 
> are {{volatile}}, although the worker is thread-confined through a 
> ThreadLocal. Removing volatile gave about 2–4%, which is within noise but 
> consistent.
> * *{{NodeHistoryAdvice}}* is always installed, with three setter calls before 
> and three after each node. It could be a single reference to a precomputed 
> immutable (id, label, source) holder.
> * *{{CamelInternalProcessor}}* calls the deprecated 
> {{notifyExchangeAsyncProcessingStartedEvent}} after every node, because the 
> error handler always returns {{false}}, without the cached 
> {{isEventNotificationApplicable()}} guard. {{SendProcessor:270}} has the same 
> issue.
> * *{{StopWatch.taken()}}* uses {{Duration.ofNanos(delta).toMillis()}}. Use 
> {{TimeUnit}}/division to avoid depending on escape analysis.
> * *{{DefaultUnitOfWork}}:* an eager {{ReentrantLock}} per UoW, taken for 
> {{pushRoute}}, {{popRoute}}, {{getRoute}} and {{routeStackLevel}}. That is 
> about 4 or more CAS per route pass. The lock exists because CAMEL-20199 
> replaced {{synchronized}}. Consider lazy synchronizations or a cheaper guard.
> * *{{RouteInflightRepositoryAdvice}}* does a {{routeCount.get(routeId)}} CHM 
> lookup twice per route pass. The LongAdder could be resolved once per route 
> start.
> * *{{TracingAdvice}} / {{BacklogTracer}} standby (JBang, dev mode)* register 
> an event notifier that turns on ExchangeSending event creation for every 
> send, even while tracing is disabled. Worth checking the cost in dev mode.
> h2. Failure path (only when failures are frequent)
> * {{ExceptionPolicy.createRedeliveryPolicy}} clones and re-parses about 24 
> options with placeholder resolution on every failed attempt.
> * {{DefaultExceptionPolicyStrategy}} allocates a TreeMap, two LinkedHashSets 
> and an ArrayList per exception, and re-splits route-scoped and context-scoped 
> policies each time.
> * The "Failed delivery …" log strings are built before deciding whether to 
> log. A DLC also regex-sanitizes {{deadLetterUri}} on every call.
> * {{RedeliveryPolicy}} uses a global lock and a shared {{Random}} for 
> collision avoidance; {{delayPattern}} is re-parsed on every redelivery.
> h2. Possible correctness bugs found along the way (need tests)
> * {{SendProcessor}} / {{DefaultProducerCache}} emit {{ExchangeSent}} only if 
> some notifier accepted {{ExchangeSending}}. {{SqlTraceDevConsole}}'s notifier 
> ignores Sending but wants Sent (CAMEL-14354 gating).
> * A nested {{Splitter}} shares or removes the failure tracker through the 
> sub-exchange ({{Splitter.afterSend}}). The outer {{SplitResult}} counts can 
> be skewed.
> * On the VIRTUAL thread-pool profile, the scheduled pool has no 
> {{removeOnCancelPolicy}}. Cancelled multicast/split timeouts keep the whole 
> {{MulticastTask}} alive until the delay expires.
> h2. Suggested order
> # Quick and low risk: #4 (text-block fast path), #2 (copy only the IN 
> message), #5 (IOConverter), the MessageSupport traits change, and 
> {{StopWatch.taken}}.
> # JMX counters: quick wins from #1(a), then the striped design.
> # Internal-properties hybrid (#3), then JMH work on the compact map.
> # Split/multicast sequential fast path and per-item allocations.
> # Throttler cleanup scheduling, the await-manager latch, and the 
> ProducerCache miss path.
> # Type-converter lookup caching (enum, try-miss) and Simple operator 
> precompilation.
> # Reactive hops and the per-node plumbing (needs design review and 
> benchmarks).
> ----
> _Claude Code on behalf of davsclaus_



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to