shashank created CAMEL-25061:
--------------------------------

             Summary: camel-file - FileLockClusterService: after the service is 
stopped and started again, the node can never become the leader again (cluster 
data tasks are rejected)
                 Key: CAMEL-25061
                 URL: https://issues.apache.org/jira/browse/CAMEL-25061
             Project: Camel
          Issue Type: Bug
          Components: camel-file
            Reporter: shashank


CAMEL-22784 (PR #20578, "Improve FileLockClusterService resilience to long 
blocking network based file I/O") runs the reads and writes of the cluster data 
file on a separate executor, so that a hung network file system cannot block 
the leadership check. {{FileLockClusterService.getClusterDataTaskExecutor()}} 
creates that executor lazily, only when the field is null. 
{{FileLockClusterService.doStop()}} shuts it down but does not reset the field, 
unlike the scheduled executor that is handled just above it, which is set to 
null after the shutdown.

When the service is stopped and started again in the same JVM, for example with 
the stop and start JMX operations of the cluster service, the views are started 
again and their leadership check runs, but {{getClusterDataTaskExecutor()}} 
returns the executor that was shut down. {{FileLockClusterTaskExecutor.run}} 
passes every task to {{CompletableFuture.supplyAsync}} on that executor, which 
throws {{RejectedExecutionException}}. The leadership check catches it, logs it 
at DEBUG level only (so nothing shows with the default logging) and tries again 
at the next interval, forever. The node can no longer read the cluster data or 
write its heartbeat, so it never becomes the leader again, and the routes it 
manages ({{master:}} routes and routes with a {{ClusteredRoutePolicy}}) never 
run on it until the JVM is restarted.

A {{stopView}}/{{startView}} of a single view is not affected, as it does not 
stop the service. A plain CamelContext stop and start does not hit it either, 
because the CamelContext shuts down and removes the services that were added to 
it, so the service is not started again with the context.

h3. Reproduction

One CamelContext with a {{FileLockClusterService}} ({{acquireLockDelay}} 100 
ms, {{acquireLockInterval}} 200 ms). Get the view of a namespace and wait until 
the node is the leader and has written its heartbeat to the data file. Call 
{{service.stop()}} (the view is stopped, the lock file is free), then 
{{service.start()}}. The view is started again, but within 10 s the node does 
not become the leader and nothing is written to the data file. With the fix 
below it is the leader again within a second, also after a second stop and 
start.

h3. Proposed fix

In {{FileLockClusterService.doStop()}}, set {{clusterDataTaskExecutor}} to null 
after shutting it down, as is already done for the scheduled executor. 
{{getClusterDataTaskExecutor()}} then creates a new one on the next start.

Affected: 4.17.0 and later, and 4.14.5 and later (the backport PR #20780 
brought the change to 4.14.x). Checked at the release tags: the executor exists 
and is not reset in 4.14.5, 4.17.0, 4.18.0 and 4.22.0, and does not exist in 
4.14.4 and 4.16.0.

Duplicate check (2026-09-28): JIRA text "FileLockClusterService" returns 5 
issues (CAMEL-25052, CAMEL-22784, CAMEL-22616, CAMEL-20997, CAMEL-11954), none 
about a restart of the service. GitHub PRs for "FileLockClusterService" are the 
CAMEL-22784 series, #20780, #24468 and #26932 (CAMEL-25052, which changes 
{{FileLockClusterView}} and does not touch this executor).

_Filed with Claude Code on behalf of allthingssecurity._




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to