Rodrigo Meneses created FLINK-40848:
---------------------------------------

             Summary: Derive FlinkDeployment resource usage from pod resource 
requests
                 Key: FLINK-40848
                 URL: https://issues.apache.org/jira/browse/FLINK-40848
             Project: Flink
          Issue Type: Improvement
          Components: Kubernetes Operator
            Reporter: Rodrigo Meneses


The FlinkDeployment status fields clusterInfo.total-cpu / total-memory (and the 
ResourceUsage.Cpu/Memory metrics built from them, documented as "requests") are 
currently computed from the Flink configuration (JM/TM resources × limit factor 
× replicas, with the TaskManager count fetched from the JobManager REST API). 
They therefore report limits rather than requests, miss sidecars and pod 
template overrides, go stale when the JobManager is unreachable or the 
deployment is suspended, and can show floating point artifacts such as 
3.3000000000000003. This proposes computing them instead as the sums of the 
resource requests of all containers of the running JobManager and TaskManager 
pods, listed from the API server cache on every observation: same fields and 
format, no new configuration, no JobManager REST call, and values that follow 
scaling, restarts and suspension.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to