Dale Lane created FLINK-40865:
---------------------------------

             Summary: Add the FlinkMiniCluster observer
                 Key: FLINK-40865
                 URL: https://issues.apache.org/jira/browse/FLINK-40865
             Project: Flink
          Issue Type: Sub-task
          Components: Kubernetes Operator
            Reporter: Dale Lane


Adds {{{}MiniClusterObserver{}}}, which will populate 
{{FlinkMiniClusterStatus}} from the observed state of a 
{{{}FlinkMiniCluster{}}}: the pod's deployment status 
({{{}status.podDeploymentStatus{}}}), cluster info, and the job status read 
from the MiniCluster's REST endpoint.

Because a MiniCluster hosts its REST endpoint in the same process as the job, 
the observer will distinguish a job reaching a terminal state from a pod 
failing, and treats an endpoint that is briefly unreachable during a pod 
restart as the expected transient it is rather than as an error (under high 
availability the launcher resubmits under the same fixed JobID and the 
dispatcher recovers the job from HA metadata).

Pod-level failures such as image pull errors and OOM kills are surfaced as 
resource errors, and stale errors are cleared once the pod is healthy again.

Nothing calls this yet, the controller wiring is a later sub-task



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to