[
https://issues.apache.org/jira/browse/FLINK-40648?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated FLINK-40648:
-----------------------------------
Labels: pull-request-available (was: )
> PrometheusReporter drops all metrics of a process when two metric group
> variables sanitise to the same label name
> -----------------------------------------------------------------------------------------------------------------
>
> Key: FLINK-40648
> URL: https://issues.apache.org/jira/browse/FLINK-40648
> Project: Flink
> Issue Type: Bug
> Components: Runtime / Metrics
> Affects Versions: 2.3.0, 2.2.1, 1.20.5, 2.4.0
> Reporter: Martijn Visser
> Assignee: Martijn Visser
> Priority: Critical
> Labels: pull-request-available
>
> AbstractPrometheusReporter.notifyOfAddedMetric builds the Prometheus label
> names by filtering every metric group variable through [^a-zA-Z0-9:_] -> _
> and adding each result to dimensionKeys, without checking whether two keys
> produced the same name.
> metricGroup.addGroup("a.b", "v1").addGroup("a-b", "v2") gives the variables
> <a.b> and <a-b>, both of which filter to a_b, and the scrape then carries
> {code}
> flink_scope_dup{a_b="v1",a_b="v2",} 7.0
> {code}
> Prometheus 3.11.2 scraping that body:
> {code}
> health = down
> lastError= label name "a_b" is not unique: invalid sample
> {code}
> The rejection is in tsdb/head_append.go, and scrape.go has no case for
> ErrInvalidSample, so it breaks out of the sample loop. The scrape is
> abandoned rather than the one sample skipped, which means every other metric
> from that TaskManager or JobManager is lost too and the target goes down.
> Verified with a body holding the colliding metric plus one valid metric: both
> come back empty and up is 0.
> Only a user-supplied addGroup key can trigger this. Flink's own variables are
> all [a-z0-9_] already.
> Two metrics named a.b and a-b in one group collide the same way one level up.
> There the second metric reuses the first one's collector, so the series shows
> whichever registered last under the first one's help string. That is the same
> defect FLINK-32649 reports, since collectorsWithCountByMetricName is keyed on
> the scoped name without the label names.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)