Claus Ibsen created CAMEL-25089:
-----------------------------------

             Summary: camel-cluster - Clustered route policy and file lock 
cluster: fix bugs found in a deep review
                 Key: CAMEL-25089
                 URL: https://issues.apache.org/jira/browse/CAMEL-25089
             Project: Camel
          Issue Type: Bug
          Components: camel-core, camel-file
            Reporter: Claus Ibsen


A review of the clustering support (camel-cluster, the cluster service support 
in camel-support, and the file lock cluster service in camel-file) found these 
bugs:

# *FileLockClusterView* does not tell its listeners that the leadership is lost 
when the view is stopped while it is the leader (the {{stopView}} JMX 
operation, stopping the cluster service, or stopping the clustered route 
controller). The clustered routes and the {{master}} consumers keep running 
while the file lock is released, so another member takes the leadership and 
both members run the routes.
# *ManagedClusterService* (JMX) {{getNamespaces}}, {{startView}}, {{stopView}} 
and {{isLeader}} look up a cluster service with the default selector instead of 
using the service of the mbean, so with more than one cluster service they do 
nothing or answer wrongly.
# *ClusteredRouteController.setClusterServiceSelector* checks the cluster 
service instead of the selector argument, so it always fails when no cluster 
service is set. The route policies it creates also ignore the selector, as the 
cluster service is not looked up yet when the policies are created.
# *ClusteredRoutePolicy*: the {{initialDelay}} is skipped when the leadership 
is taken after the CamelContext has started (which is common with the file lock 
cluster service), and the routes are started right away.
# *ClusteredRoutePolicy*: a removed route is kept in the started/stopped routes 
of a shared policy, so the policy stops (or starts) another route that is added 
later with the same route id when the leadership changes.
# *AbstractCamelClusterService.doStart* starts all its views, also the views 
that are no longer used (their routes have been removed), so after a stop/start 
of the cluster service a file lock member takes a lock of a namespace it does 
not use.
# *FileLockClusterView* fails with ArithmeticException (/ by zero) when the 
{{acquireLockInterval}} is less than 1 millisecond (such as 500 MICROSECONDS). 
It is now rejected with a clear error.
# The cluster services found by type are returned in no particular order, so 
the {{first}} (and type/attribute) selectors do not pick the same service every 
time.

Not changed:
* The route policy created per route by ClusteredRoutePolicyFactory / 
ClusteredRouteController is not cleaned up (its event notifier, thread pool and 
service) when the route is removed, only when the CamelContext is shut down. A 
shared policy can be used again by a route added later, so this needs its own 
change.
* After a stop/start of the CamelContext the clustered routes are not started 
again (the policy thread pool is shut down on stop).

_Claude Code on behalf of Claus Ibsen_



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to