Claus Ibsen created CAMEL-25089:
-----------------------------------
Summary: camel-cluster - Clustered route policy and file lock
cluster: fix bugs found in a deep review
Key: CAMEL-25089
URL: https://issues.apache.org/jira/browse/CAMEL-25089
Project: Camel
Issue Type: Bug
Components: camel-core, camel-file
Reporter: Claus Ibsen
A review of the clustering support (camel-cluster, the cluster service support
in camel-support, and the file lock cluster service in camel-file) found these
bugs:
# *FileLockClusterView* does not tell its listeners that the leadership is lost
when the view is stopped while it is the leader (the {{stopView}} JMX
operation, stopping the cluster service, or stopping the clustered route
controller). The clustered routes and the {{master}} consumers keep running
while the file lock is released, so another member takes the leadership and
both members run the routes.
# *ManagedClusterService* (JMX) {{getNamespaces}}, {{startView}}, {{stopView}}
and {{isLeader}} look up a cluster service with the default selector instead of
using the service of the mbean, so with more than one cluster service they do
nothing or answer wrongly.
# *ClusteredRouteController.setClusterServiceSelector* checks the cluster
service instead of the selector argument, so it always fails when no cluster
service is set. The route policies it creates also ignore the selector, as the
cluster service is not looked up yet when the policies are created.
# *ClusteredRoutePolicy*: the {{initialDelay}} is skipped when the leadership
is taken after the CamelContext has started (which is common with the file lock
cluster service), and the routes are started right away.
# *ClusteredRoutePolicy*: a removed route is kept in the started/stopped routes
of a shared policy, so the policy stops (or starts) another route that is added
later with the same route id when the leadership changes.
# *AbstractCamelClusterService.doStart* starts all its views, also the views
that are no longer used (their routes have been removed), so after a stop/start
of the cluster service a file lock member takes a lock of a namespace it does
not use.
# *FileLockClusterView* fails with ArithmeticException (/ by zero) when the
{{acquireLockInterval}} is less than 1 millisecond (such as 500 MICROSECONDS).
It is now rejected with a clear error.
# The cluster services found by type are returned in no particular order, so
the {{first}} (and type/attribute) selectors do not pick the same service every
time.
Not changed:
* The route policy created per route by ClusteredRoutePolicyFactory /
ClusteredRouteController is not cleaned up (its event notifier, thread pool and
service) when the route is removed, only when the CamelContext is shut down. A
shared policy can be used again by a route added later, so this needs its own
change.
* After a stop/start of the CamelContext the clustered routes are not started
again (the policy thread pool is shut down on stop).
_Claude Code on behalf of Claus Ibsen_
--
This message was sent by Atlassian Jira
(v8.20.10#820010)