shashank created CAMEL-25209:
--------------------------------
Summary: camel-kubernetes - a failed Lease renewal stops the
KubernetesLeadershipController for good: the leader gives up its routes and no
other pod takes over
Key: CAMEL-25209
URL: https://issues.apache.org/jira/browse/CAMEL-25209
Project: Camel
Issue Type: Bug
Components: camel-kubernetes
Reporter: shashank
{{KubernetesLeadershipController}} runs the leader election as a chain of tasks
on a single-thread scheduled executor: each {{refreshStatus()}} run schedules
the next one ({{rescheduleAfterDelay()}} or
{{serializedExecutor.execute(this::refreshStatus)}}). Every Kubernetes call in
it catches its exceptions and reschedules ({{lookupNewLeaderInfo}},
{{tryAcquireLeadership}}, {{yieldLeadership}}), except the Lease renewal in the
LEADER state:
{code:java}
HasMetadata newLease =
this.leaseManager.refreshLeaseRenewTime(kubernetesClient,
this.latestLeaseResource,
this.lockConfiguration.getRenewDeadlineSeconds());
updateLatestLeaderInfo(newLease, this.latestMembers);
rescheduleAfterDelay();
{code}
With the default resource type {{Lease}},
{{NativeLeaseResourceManager.refreshLeaseRenewTime}} does a PUT with
{{lockResourceVersion}} once per renew deadline. When that PUT fails, the
{{KubernetesClientException}} leaves {{refreshStatus}}, the executor keeps it
in the (unread) future, and the next run is never scheduled. Nothing is logged.
The Kubernetes client retries a 429, a 5xx or an I/O error itself (by default
up to 10 times with an exponential backoff, about 20 s), so the PUT fails when
the API server cannot be reached or answers with errors for longer than that
(an API server restart, a control plane upgrade, a network problem), or with an
error that is not retried (a 409 conflict, 401/403, 404 when the Lease was
deleted).
Then:
* on the leader pod, the {{TimedLeaderNotifier}} is no longer refreshed, so
after the renew deadline it reports "no leader": the pod stops its master
routes / clustered routes;
* the other pods keep reading a Lease held by a pod that is running and ready,
which {{LeaderInfo.hasValidLeader()}} accepts (it does not look at
{{renewTime}}), so none of them tries to acquire it.
No pod of the group leads until the former leader pod is restarted (or becomes
not ready). One failed renewal is enough; a failed lookup of the Lease during
the same outage is harmless, as it is caught and retried.
The {{ConfigMap}} resource type is not affected ({{refreshLeaseRenewTime}} does
nothing there). The unguarded call came with the Lease support in CAMEL-15881
(2020); the other calls were guarded from the start.
h3. Reproduction
{{KubernetesClusterServiceTest}} with the existing {{LockTestServer}} (mock API
server), Lease type, two pods. When a leader is elected, the leader's server
refuses PUT requests on the Lease (500) until one renewal has been refused,
then accepts them again (the test client has the client retries disabled, as in
the other tests of the module). Without the fix the leader's recorder reports
{{null}} after the renew deadline and stays so ({{expected: <mypod1> but was:
<null>}}, three runs), and the test log has no warning about the failed renewal.
h3. Proposed fix
{{refreshStatus()}} wraps the state machine in a {{try/catch}}: on an exception
it logs a WARN (stack trace at DEBUG) and schedules the next run after the
retry period, like the other failure paths; after {{stop()}} it only logs at
DEBUG. In the LEADER state the next run reads the Lease again and renews it.
Test:
{{KubernetesClusterServiceTest.testLeaderKeepsLeadershipAfterFailedLeaseRenewal}},
with a new option of {{LockTestServer}} to refuse only the update (PUT)
requests. It fails without the fix.
Affected: all versions with the Lease resource type (the same unguarded call at
camel-3.20.0, 4.18.0 and main).
Duplicate check (2026-09-30): JIRA text "KubernetesLeadershipController",
"refreshLeaseRenewTime", "kubernetes leadership renew", and component
camel-kubernetes with leader/leadership/cluster/lease in the summary: only
CAMEL-21720, CAMEL-15881, CAMEL-19825, CAMEL-11837 (different problems). GitHub
pull requests "KubernetesLeadershipController", "kubernetes leadership": none.
No open pull request touches these files.
_Filed with Claude Code on behalf of allthingssecurity._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)