[
https://issues.apache.org/jira/browse/CAMEL-25052?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Claus Ibsen resolved CAMEL-25052.
---------------------------------
Resolution: Fixed
Fixed by https://github.com/apache/camel/pull/26932 (merged to main for 4.23.0).
_Claude Code on behalf of davsclaus_
> camel-file - FileLockClusterView: stopping the view no longer ends a
> leadership check that is already running (regression from CAMEL-22784), so a
> stopped view can take the cluster lock, and a quick restart doubles the checks
> --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
>
> Key: CAMEL-25052
> URL: https://issues.apache.org/jira/browse/CAMEL-25052
> Project: Camel
> Issue Type: Bug
> Components: camel-file
> Reporter: shashank
> Priority: Minor
> Fix For: 4.23.0
>
>
> Since CAMEL-22784 ("Use predictable scheduling for FileLockClusterService
> lock acquisition", PR #20686, commit 7ee5afcd2f), {{FileLockClusterView}}
> runs its leadership check as a chain of one-shot tasks: {{scheduleTryLock}}
> ({{FileLockClusterView.java:296-324}}) calls
> {{executor.schedule(this::tryLock, ...)}}, and {{tryLock}} schedules the next
> run in its {{finally}} block ({{:280}}). The {{ScheduledFuture}} is no longer
> stored in the {{task}} field ({{:60}}). {{task}} therefore stays null, and
> {{closeInternal()}} ({{:156-163}}), which {{doStop}} ({{:133-154}}) calls,
> cancels nothing. Before that commit, {{task.cancel(true)}} interrupted a
> running check.
> {{tryLock}} checks {{isStarting() || isStarted()}} only once, when it begins
> ({{:196}}). This has two consequences.
> # *A stop during a running check.* A follower that sees a stale or absent
> leader opens the lock and data files and calls {{FileChannel.tryLock}}
> ({{:250-263}}). If the view is stopped meanwhile, the check carries on after
> {{doStop}} has closed the files. It takes the file lock, sets the member to
> LEADER, fires the leadership event and writes its heartbeat once, on a view
> that is Stopped. The next run of the chain sees the view stopped and ends.
> The stopped node reports {{isLeader(namespace)=true}}, and the OS lock stays
> held until the view is started again ({{doStart}} closes leftover files at
> {{:101-104}}) or the JVM exits. Meanwhile no other node can become the
> leader, so the clustered routes run nowhere.
> # *A stop and start within one {{acquireLockInterval}}.* The chain of the
> previous start is still scheduled, sees the view started again and keeps
> running next to the chain of the new start. Every quick restart adds one more
> chain.
> The first case needs the view to be stopped while its node is inside the
> acquisition I/O (it has just seen the old leader go away), and the JVM to
> keep running afterwards. The view is stopped when its last user releases it:
> a camel-master route stops, the last route with a {{ClusteredRoutePolicy}} is
> removed, or the CamelContext stops in a JVM that keeps running (an
> application server, a Spring context refresh). JMX {{stopView}} also stops
> it. All of these can coincide with a leadership handover. The window is the
> acquisition I/O: small on a local disk, larger on the network storage this
> service targets ({{clusterDataTaskTimeout}} defaults to 10 s, with 5 attempts
> per task).
> h3. Reproduction
> Each node is a CamelContext with its own {{FileLockClusterService}} on a
> shared root ({{acquireLockInterval}} 500 ms) and a {{master:}} route.
> * A is the leader. B follows. A's context stops, and B's next check is held
> (test hook in {{createRandomAccessFile}}) just before it opens the lock file.
> B's master route is stopped, which stops B's view, then the hold is released.
> Result: B's view is Stopped but {{isLeader(ns)=true}}. A probe from a
> separate JVM finds the lock held, and a node C in a separate JVM does not
> become the leader within 6 s.
> * The unit test in the PR does the same in one JVM, with B's view released
> directly: B reports leadership and the lock file stays locked by B's channel.
> * Without a hook, with 200 ms file opens (as on network storage) and B's
> master route stopped at a random time during the takeover: 3 of 15 rounds
> ended with the stopped view holding the lock. With local file opens: 0 of 15.
> * One node, three quick {{stopRoute}}/{{startRoute}} of its master route: the
> leadership check runs 10 times per 5 s before and 40 times per 5 s after.
> A TLA+ model of two or three nodes (the check, view stop and start, the OS
> lock, and the camel-master listener) finds the same traces:
> {{NoStoppedHolder}} and {{StoppedNotLeader}} are violated by the trace
> TickAcquire(a), ConsumerStop(a), Acquire(a); {{EventuallyLeader}} is
> violated; {{OneChain}} is violated by a quick stop and start. The fixed model
> holds all properties.
> h3. Proposed fix
> * Give every start a generation number, incremented on every start and stop
> under a small state lock. A check ends without rescheduling itself once its
> generation is over, which ends the chain on stop and keeps one chain per
> start. This replaces the dead {{task}} field.
> * Open the files and take the OS lock into local variables, without holding
> any lock. Then re-check the generation under the state lock and only then
> publish the lock and files to the view and set LEADER. If the view was
> stopped meanwhile, release the lock and close the files instead.
> * {{doStop}} increments the generation and takes over the lock and files
> under the same state lock, then does the file I/O (truncate, release, close)
> outside it.
> * The leadership event is fired outside the state lock, as before. The state
> lock is never held during file I/O or listener callbacks, so a stop does not
> wait for a slow (NFS) check.
> Affected: 4.17.0 and later, and 4.14.5 and later (the backport PR #20780
> brought the change to 4.14.x). Checked at the release tags: the {{task =}}
> assignment is present in 4.16.0 and 4.14.4, and absent in 4.14.5, 4.17.0,
> 4.18.0, 4.18.4, 4.22.0 and main.
> Duplicate check (2026-09-27): JIRA text "FileLockClusterView" (6 issues),
> "FileLockClusterService" (4) and "file lock cluster" (15) return only the
> CAMEL-22430 and CAMEL-22784 work and CAMEL-22541 (a flaky test). GitHub PRs
> for these names are the CAMEL-22784 series (#20433, #20452, #20526, #20578,
> #20590, #20686) and #20780. Nothing covers the stop.
> _Filed with Claude Code on behalf of allthingssecurity._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)