shashank created CAMEL-25061:
--------------------------------
Summary: camel-file - FileLockClusterService: after the service is
stopped and started again, the node can never become the leader again (cluster
data tasks are rejected)
Key: CAMEL-25061
URL: https://issues.apache.org/jira/browse/CAMEL-25061
Project: Camel
Issue Type: Bug
Components: camel-file
Reporter: shashank
CAMEL-22784 (PR #20578, "Improve FileLockClusterService resilience to long
blocking network based file I/O") runs the reads and writes of the cluster data
file on a separate executor, so that a hung network file system cannot block
the leadership check. {{FileLockClusterService.getClusterDataTaskExecutor()}}
creates that executor lazily, only when the field is null.
{{FileLockClusterService.doStop()}} shuts it down but does not reset the field,
unlike the scheduled executor that is handled just above it, which is set to
null after the shutdown.
When the service is stopped and started again in the same JVM, for example with
the stop and start JMX operations of the cluster service, the views are started
again and their leadership check runs, but {{getClusterDataTaskExecutor()}}
returns the executor that was shut down. {{FileLockClusterTaskExecutor.run}}
passes every task to {{CompletableFuture.supplyAsync}} on that executor, which
throws {{RejectedExecutionException}}. The leadership check catches it, logs it
at DEBUG level only (so nothing shows with the default logging) and tries again
at the next interval, forever. The node can no longer read the cluster data or
write its heartbeat, so it never becomes the leader again, and the routes it
manages ({{master:}} routes and routes with a {{ClusteredRoutePolicy}}) never
run on it until the JVM is restarted.
A {{stopView}}/{{startView}} of a single view is not affected, as it does not
stop the service. A plain CamelContext stop and start does not hit it either,
because the CamelContext shuts down and removes the services that were added to
it, so the service is not started again with the context.
h3. Reproduction
One CamelContext with a {{FileLockClusterService}} ({{acquireLockDelay}} 100
ms, {{acquireLockInterval}} 200 ms). Get the view of a namespace and wait until
the node is the leader and has written its heartbeat to the data file. Call
{{service.stop()}} (the view is stopped, the lock file is free), then
{{service.start()}}. The view is started again, but within 10 s the node does
not become the leader and nothing is written to the data file. With the fix
below it is the leader again within a second, also after a second stop and
start.
h3. Proposed fix
In {{FileLockClusterService.doStop()}}, set {{clusterDataTaskExecutor}} to null
after shutting it down, as is already done for the scheduled executor.
{{getClusterDataTaskExecutor()}} then creates a new one on the next start.
Affected: 4.17.0 and later, and 4.14.5 and later (the backport PR #20780
brought the change to 4.14.x). Checked at the release tags: the executor exists
and is not reset in 4.14.5, 4.17.0, 4.18.0 and 4.22.0, and does not exist in
4.14.4 and 4.16.0.
Duplicate check (2026-09-28): JIRA text "FileLockClusterService" returns 5
issues (CAMEL-25052, CAMEL-22784, CAMEL-22616, CAMEL-20997, CAMEL-11954), none
about a restart of the service. GitHub PRs for "FileLockClusterService" are the
CAMEL-22784 series, #20780, #24468 and #26932 (CAMEL-25052, which changes
{{FileLockClusterView}} and does not touch this executor).
_Filed with Claude Code on behalf of allthingssecurity._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)