Rodrigo Meneses created FLINK-40848:
---------------------------------------
Summary: Derive FlinkDeployment resource usage from pod resource
requests
Key: FLINK-40848
URL: https://issues.apache.org/jira/browse/FLINK-40848
Project: Flink
Issue Type: Improvement
Components: Kubernetes Operator
Reporter: Rodrigo Meneses
The FlinkDeployment status fields clusterInfo.total-cpu / total-memory (and the
ResourceUsage.Cpu/Memory metrics built from them, documented as "requests") are
currently computed from the Flink configuration (JM/TM resources × limit factor
× replicas, with the TaskManager count fetched from the JobManager REST API).
They therefore report limits rather than requests, miss sidecars and pod
template overrides, go stale when the JobManager is unreachable or the
deployment is suspended, and can show floating point artifacts such as
3.3000000000000003. This proposes computing them instead as the sums of the
resource requests of all containers of the running JobManager and TaskManager
pods, listed from the API server cache on every observation: same fields and
format, no new configuration, no JobManager REST call, and values that follow
scaling, restarts and suspension.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)