Dale Lane created FLINK-40865:
---------------------------------
Summary: Add the FlinkMiniCluster observer
Key: FLINK-40865
URL: https://issues.apache.org/jira/browse/FLINK-40865
Project: Flink
Issue Type: Sub-task
Components: Kubernetes Operator
Reporter: Dale Lane
Adds {{{}MiniClusterObserver{}}}, which will populate
{{FlinkMiniClusterStatus}} from the observed state of a
{{{}FlinkMiniCluster{}}}: the pod's deployment status
({{{}status.podDeploymentStatus{}}}), cluster info, and the job status read
from the MiniCluster's REST endpoint.
Because a MiniCluster hosts its REST endpoint in the same process as the job,
the observer will distinguish a job reaching a terminal state from a pod
failing, and treats an endpoint that is briefly unreachable during a pod
restart as the expected transient it is rather than as an error (under high
availability the launcher resubmits under the same fixed JobID and the
dispatcher recovers the job from HA metadata).
Pod-level failures such as image pull errors and OOM kills are surfaced as
resource errors, and stale errors are cleared once the pod is healthy again.
Nothing calls this yet, the controller wiring is a later sub-task
--
This message was sent by Atlassian Jira
(v8.20.10#820010)