[ 
https://issues.apache.org/jira/browse/CAMEL-25209?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18121459#comment-18121459
 ] 

Claus Ibsen commented on CAMEL-25209:
-------------------------------------

Merged in https://github.com/apache/camel/pull/27158 (commit 97e434aa597d), 
thanks!

_Claude Code on behalf of davsclaus_

> camel-kubernetes - a failed Lease renewal stops the 
> KubernetesLeadershipController for good: the leader gives up its routes and 
> no other pod takes over
> -------------------------------------------------------------------------------------------------------------------------------------------------------
>
>                 Key: CAMEL-25209
>                 URL: https://issues.apache.org/jira/browse/CAMEL-25209
>             Project: Camel
>          Issue Type: Bug
>          Components: camel-kubernetes
>            Reporter: shashank
>            Assignee: shashank
>            Priority: Major
>             Fix For: 4.23.0
>
>
> {{KubernetesLeadershipController}} runs the leader election as a chain of 
> tasks on a single-thread scheduled executor: each {{refreshStatus()}} run 
> schedules the next one ({{rescheduleAfterDelay()}} or 
> {{serializedExecutor.execute(this::refreshStatus)}}). Every Kubernetes call 
> in it catches its exceptions and reschedules ({{lookupNewLeaderInfo}}, 
> {{tryAcquireLeadership}}, {{yieldLeadership}}), except the Lease renewal in 
> the LEADER state:
> {code:java}
> HasMetadata newLease = 
> this.leaseManager.refreshLeaseRenewTime(kubernetesClient, 
> this.latestLeaseResource,
>         this.lockConfiguration.getRenewDeadlineSeconds());
> updateLatestLeaderInfo(newLease, this.latestMembers);
> rescheduleAfterDelay();
> {code}
> With the default resource type {{Lease}}, 
> {{NativeLeaseResourceManager.refreshLeaseRenewTime}} does a PUT with 
> {{lockResourceVersion}} once per renew deadline. When that PUT fails, the 
> {{KubernetesClientException}} leaves {{refreshStatus}}, the executor keeps it 
> in the (unread) future, and the next run is never scheduled. Nothing is 
> logged. The Kubernetes client retries a 429, a 5xx or an I/O error itself (by 
> default up to 10 times with an exponential backoff, about 20 s), so the PUT 
> fails when the API server cannot be reached or answers with errors for longer 
> than that (an API server restart, a control plane upgrade, a network 
> problem), or with an error that is not retried (a 409 conflict, 401/403, 404 
> when the Lease was deleted).
> Then:
> * on the leader pod, the {{TimedLeaderNotifier}} is no longer refreshed, so 
> after the renew deadline it reports "no leader": the pod stops its master 
> routes / clustered routes;
> * the other pods keep reading a Lease held by a pod that is running and 
> ready, which {{LeaderInfo.hasValidLeader()}} accepts (it does not look at 
> {{renewTime}}), so none of them tries to acquire it.
> No pod of the group leads until the former leader pod is restarted (or 
> becomes not ready). One failed renewal is enough; a failed lookup of the 
> Lease during the same outage is harmless, as it is caught and retried.
> The {{ConfigMap}} resource type is not affected ({{refreshLeaseRenewTime}} 
> does nothing there). The unguarded call came with the Lease support in 
> CAMEL-15881 (2020); the other calls were guarded from the start.
> h3. Reproduction
> {{KubernetesClusterServiceTest}} with the existing {{LockTestServer}} (mock 
> API server), Lease type, two pods. When a leader is elected, the leader's 
> server refuses PUT requests on the Lease (500) until one renewal has been 
> refused, then accepts them again (the test client has the client retries 
> disabled, as in the other tests of the module). Without the fix the leader's 
> recorder reports {{null}} after the renew deadline and stays so ({{expected: 
> <mypod1> but was: <null>}}, three runs), and the test log has no warning 
> about the failed renewal.
> h3. Proposed fix
> {{refreshStatus()}} wraps the state machine in a {{try/catch}}: on an 
> exception it logs a WARN (stack trace at DEBUG) and schedules the next run 
> after the retry period, like the other failure paths; after {{stop()}} it 
> only logs at DEBUG. In the LEADER state the next run reads the Lease again 
> and renews it.
> Test: 
> {{KubernetesClusterServiceTest.testLeaderKeepsLeadershipAfterFailedLeaseRenewal}},
>  with a new option of {{LockTestServer}} to refuse only the update (PUT) 
> requests. It fails without the fix.
> Affected: all versions with the Lease resource type (the same unguarded call 
> at camel-3.20.0, 4.18.0 and main).
> Duplicate check (2026-09-30): JIRA text "KubernetesLeadershipController", 
> "refreshLeaseRenewTime", "kubernetes leadership renew", and component 
> camel-kubernetes with leader/leadership/cluster/lease in the summary: only 
> CAMEL-21720, CAMEL-15881, CAMEL-19825, CAMEL-11837 (different problems). 
> GitHub pull requests "KubernetesLeadershipController", "kubernetes 
> leadership": none. No open pull request touches these files.
> _Filed with Claude Code on behalf of allthingssecurity._



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to