[
https://issues.apache.org/jira/browse/HDDS-16611?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Chi-Hsuan Huang updated HDDS-16611:
-----------------------------------
Description:
{{OMMetrics}} declares {{numVolumes}}, {{numBuckets}} and {{numKeys}} as
{{MutableCounterLong}}, but they behave like gauges. They are decremented with
{{incr(-1)}}, and {{setNum*()}} emulates a set via {{incr(val - oldVal)}}. This
has been the case since HDDS-816.
The JMX values are correct, but {{/prom}} exports these metrics as {{# TYPE
counter}}. Prometheus treats any decrease in a counter as a reset. So the *Key
Creation Rate* panel in the Memory Consumption dashboard, which uses
{{rate(om_metrics_num_keys[1m])}}, reports inflated values whenever keys are
deleted.
HDDS-10597 fixed the same pattern in {{SafeModeMetrics}}.
Proposed changes:
* Switch the three fields to {{MutableGaugeLong}} and use {{set()}} in
{{setNumVolumes}}, {{setNumBuckets}} and {{setNumKeys}}.
* Update {{TestOmMetrics}} to read these values with {{getLongGauge}} instead
of {{getLongCounter}}.
* Update the Key Creation Rate panel to use {{deriv()}}, or
{{rate(om_metrics_num_key_commits[1m])}}.
Metric names in {{/jmx}} and {{/prom}} stay the same.
was:
{{OMMetrics}} declares {{numVolumes}}, {{numBuckets}} and {{numKeys}} as
{{MutableCounterLong}}, but they behave like gauges. They are decremented with
{{incr\(\-1\)}}, and {{setNum\*\(\)}} emulates a set via {{incr\(val \-
oldVal\)}}. This has been the case since HDDS\-816.
The JMX values are correct, but {{/prom}} exports these metrics as {{# TYPE
counter}}. Prometheus treats any decrease in a counter as a reset. So the _Key
Creation Rate_ panel in the Memory Consumption dashboard, which uses
{{rate\(om\_metrics\_num\_keys\[1m\]\)}}, reports inflated values whenever keys
are deleted.
HDDS\-10597 fixed the same pattern in {{SafeModeMetrics}}.
Proposed changes:
\- Switch the three fields to {{MutableGaugeLong}} and use {{set\(\)}} in
{{setNumVolumes}}, {{setNumBuckets}} and {{setNumKeys}}.
\- Update {{TestOmMetrics}} to read these values with {{getLongGauge}} instead
of {{getLongCounter}}.
\- Update the Key Creation Rate panel to use {{deriv\(\)}}, or
{{rate\(om\_metrics\_num\_key\_commits\[1m\]\)}}.
Metric names in {{/jmx}} and {{/prom}} stay the same.
> Use MutableGaugeLong for numVolumes, numBuckets and numKeys in OMMetrics
> ------------------------------------------------------------------------
>
> Key: HDDS-16611
> URL: https://issues.apache.org/jira/browse/HDDS-16611
> Project: Apache Ozone
> Issue Type: Bug
> Components: Ozone Manager
> Reporter: Chi-Hsuan Huang
> Assignee: Chi-Hsuan Huang
> Priority: Major
>
> {{OMMetrics}} declares {{numVolumes}}, {{numBuckets}} and {{numKeys}} as
> {{MutableCounterLong}}, but they behave like gauges. They are decremented
> with {{incr(-1)}}, and {{setNum*()}} emulates a set via {{incr(val -
> oldVal)}}. This has been the case since HDDS-816.
> The JMX values are correct, but {{/prom}} exports these metrics as {{# TYPE
> counter}}. Prometheus treats any decrease in a counter as a reset. So the
> *Key Creation Rate* panel in the Memory Consumption dashboard, which uses
> {{rate(om_metrics_num_keys[1m])}}, reports inflated values whenever keys are
> deleted.
> HDDS-10597 fixed the same pattern in {{SafeModeMetrics}}.
> Proposed changes:
> * Switch the three fields to {{MutableGaugeLong}} and use {{set()}} in
> {{setNumVolumes}}, {{setNumBuckets}} and {{setNumKeys}}.
> * Update {{TestOmMetrics}} to read these values with {{getLongGauge}} instead
> of {{getLongCounter}}.
> * Update the Key Creation Rate panel to use {{deriv()}}, or
> {{rate(om_metrics_num_key_commits[1m])}}.
> Metric names in {{/jmx}} and {{/prom}} stay the same.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]