hyeonkimmm opened a new pull request, #1216: URL: https://github.com/apache/flink-kubernetes-operator/pull/1216
## What is the purpose of the change Fixes [FLINK-40470](https://issues.apache.org/jira/browse/FLINK-40470): a savepoint redeploy (`job.savepointRedeployNonce` change) whose deployment attempt fails loops forever. `AbstractJobReconciler#redeployWithSavepoint` was the only deployment path that recorded the target spec in the status *after* the `deploy` call. When `deploy` threw (transient API server error, or an error raised after the cluster resources were already created, e.g. resolving the rest endpoint of a freshly created LoadBalancer service on an RBAC-restricted cluster), the new nonce never reached `lastReconciledSpec`. Every subsequent reconciliation classified the diff as `SAVEPOINT_REDEPLOY` again and repeated the full cancel + redeploy cycle until manual intervention. As discussed on the ticket, this PR also stops the resume path from discarding the `initialSavepointPath` of a savepoint redeploy when the spec declares `upgradeMode: stateless`. Without it a savepoint redeploy requested while suspended started from empty state on resume, and the retry introduced here would have done the same for stateless jobs. ## Brief change log - `redeployWithSavepoint` records the target spec with `updateStatusBeforeDeploymentAttempt` + `patchAndCacheStatus` before calling `deploy`, like the suspend/resume and first deployment paths. A failed attempt is retried through the regular restore path from the recorded `upgradeSavepointPath` instead of triggering a new redeploy. - The suspended -> running path keeps inheriting the recorded upgrade mode for stateless specs when the recorded `upgradeSavepointPath` is the savepoint the user explicitly requested through `initialSavepointPath` (`restoringFromInitialSavepoint`). Resuming a job suspended in `savepoint` or `last-state` mode with a stateless spec still starts from empty state, as `ApplicationReconcilerUpgradeModeTest#testUpgradeJmDeployCannotStart` expects; I kept that escape hatch on purpose, happy to drop it if you prefer to never discard `upgradeSavepointPath` for stateless specs (it is a one-line change). - Docs: noted in *Suspending and Resuming* that a savepoint redeploy requested while suspended is honoured on resume regardless of the upgrade mode. ## Verifying this change This change added tests and can be verified as follows: - `ApplicationReconcilerTest#testSavepointRedeployRecoversFromFailedDeployment` and `SessionJobReconcilerTest#testSavepointRedeployRecoversFromFailedDeployment`: the deploy of a savepoint redeploy fails; the nonce, `SUSPENDED` state and `SAVEPOINT` mode must already be recorded in `lastReconciledSpec` with the status in `UPGRADING`, and the next reconciliation must restore from the recorded savepoint. Fails before the fix (status stays `DEPLOYED`, nonce unrecorded). - `ApplicationReconcilerTest#testResumeAfterSavepointRedeployWhileSuspended` (all upgrade modes): a savepoint redeploy requested while suspended must be honoured on resume. Fails before the fix for `STATELESS` (job started from empty state). - `mvn -pl flink-kubernetes-operator -am test` passes locally. ## Does this pull request potentially affect one of the following parts: - Dependencies (does it add or upgrade a dependency): no - The public API, i.e., is any changes to the `CustomResourceDescriptors`: no - Core observer or reconciler logic that is regularly executed: yes ## Documentation - Does this pull request introduce a new feature? no - If yes, how is the feature documented? not applicable --- ##### Was generative AI tooling used to co-author this PR? - [X] Yes (Claude Code) Generated-by: Claude Code (Claude Fable 5.1) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
