Claus Ibsen created CAMEL-25268:
-----------------------------------

             Summary: camel-cli-connector - the file transport should not let a 
slow status hold up actions
                 Key: CAMEL-25268
                 URL: https://issues.apache.org/jira/browse/CAMEL-25268
             Project: Camel
          Issue Type: Improvement
          Components: camel-jbang
            Reporter: Claus Ibsen


The file transport ({{FileCliConnectorTransport}}) does all its work on a 
single thread: every second (100 ms when debugging) {{task()}} polls for action 
files, then collects the status, then (every second round) the trace, debug, 
history, error, receive and activity snapshots.

On a large integration the status can take tens of seconds to collect (316 YAML 
routes / ~4000 processors took about 30 s, see CAMEL-25267). While it runs:

* an action file is only picked up when the next {{task()}} starts, so it can 
wait up to the full collection time. The CLI waits 10 s for the output file by 
default ({{ActionBaseCommand.getJsonObject}}), so {{camel cmd ...}} actions 
against a large integration can time out intermittently.
* deleting the lock file ({{camel stop}}) is only noticed when the next 
{{task()}} starts.
* the debug snapshot (breakpoints) waits behind the status.
* with a fixed 1 s delay after a 30 s collection, one CPU core is busy almost 
all the time.

The WebSocket transport had the same problem, shown as dropped connections (the 
heartbeat was starved), and was fixed in the CAMEL-25197 follow-up (PR 
https://github.com/apache/camel/pull/27279) by collecting snapshots on their 
own thread and waiting at least as long as the last collection took.

Possible changes for the file transport:
# poll for actions on a separate thread from snapshot collection (actions are 
cheap and need low latency; the dispatcher is not thread-safe, so actions 
should still run one at a time)
# schedule the next status round after max(delay, time the last round took)

CAMEL-25267 (making the status cheap to collect) fixes the cause for both 
transports; this issue is about the file transport staying responsive when 
collection is slow anyway.

_Claude Code on behalf of davsclaus_



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to