Hi all,
We hit a production incident where a JobManager that lost a leader election
deleted the HA ConfigMap belonging to a job that a different JobManager had
just taken over. I would like to know whether there is a configuration that
prevents this, and whether the cleanup path itself is something that should
be fixed in code.
*Setup*
- Flink 2.2.1 (Rev:450c63e), application mode, Kubernetes HA
- Flink Kubernetes Operator 1.15.0
- jobManager.replicas: 2
- upgradeMode: last-state
*What happened*
The operator performed a routine last-state upgrade. It deleted and
recreated the JobManager Deployment. Both new JobManager pods entered the
leader election in the same second. Pod A won, then lost the lease to pod B
about 26 seconds later. On revocation, pod A ran job-scoped HA cleanup and
deleted <clusterId>-<jobId>-config-map, which pod B was actively using.
14:27:18 [operator] Completed Deleting JobManager Deployment
14:27:18 [operator] Keeping HA metadata for last-state restore
14:27:48 [pod A] DefaultDispatcherRunner was granted leadership with
leader id 97c1e409-8f78-4a4b-b340-7fee58502285.
14:27:52 [pod A] EmbeddedExecutor - Submitting Job with JobId=62819cb2...
14:28:14 [pod B] DefaultDispatcherRunner was granted leadership with
leader id 4e0a80e3-d4f3-466c-b268-16865b2f1507.
14:28:22 [pod B] JobMasterServiceLeadershipRunner for job 62819cb2... was
granted leadership with leader id 4e0a80e3-...
14:28:22 [pod B] Found 3 checkpoints in
KubernetesStateHandleStore{configMapName='<cluster>-62819cb2...-config-map'}
14:28:25 [pod A] DefaultDispatcherRunner was revoked the leadership with
leader id 97c1e409-... Stopping the DispatcherLeaderProcess.
14:28:25 [pod A] KubernetesLeaderElectionHaServices - Clean up the high
availability data for job 62819cb20caec54fb13b0cbab17e4a9d.
14:28:25 [pod A] StandaloneDispatcher - Stopping all currently running
jobs of dispatcher
>From 14:33 onward, every checkpoint trigger on pod B failed, once per
5-minute interval (our checkpoint interval):
CheckpointFailureManager - Failed to trigger or complete checkpoint
UNKNOWN_CHECKPOINT_ID for job 62819cb2... (0 consecutive failed attempts so
far)
| error.message: Trigger checkpoint failure
When the leading JobManager was later restarted, the cluster could not
bring the job back:
Recover all persisted job graphs that are not finished, yet.
Retrieved job ids [62819cb2...] from KubernetesStateHandleStore{
configMapName='<cluster>-cluster-config-map'}
Recovered StreamGraph(jobId: 62819cb2...).
Successfully recovered 1 persisted job graphs.
ClusterEntrypoint - Fatal error occurred in the cluster entrypoint.
| error.message: Could not start recovered job
62819cb20caec54fb13b0cbab17e4a9d.
The cluster ConfigMap still listed the job id, but its per-job ConfigMap
was gone. The JobManager crashed, restarted, recovered the same job graph,
and crashed again, for roughly 20 minutes. It only stopped when we forced a
new job id via restart from the latest checkpoint.
If this looks like a genuine bug, I am happy to file a Jira.
Thanks,
Piotr