[ 
https://issues.apache.org/jira/browse/SPARK-59949?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

ASF GitHub Bot updated SPARK-59949:
-----------------------------------
    Labels: pull-request-available  (was: )

> Record the `operator.sdk` controller execution histograms with sub-second 
> precision
> -----------------------------------------------------------------------------------
>
>                 Key: SPARK-59949
>                 URL: https://issues.apache.org/jira/browse/SPARK-59949
>             Project: Spark
>          Issue Type: Sub-task
>          Components: Kubernetes
>    Affects Versions: kubernetes-operator-1.0.0
>            Reporter: Peter Toth
>            Priority: Major
>              Labels: pull-request-available
>
> - {{OperatorJosdkMetrics.timeControllerExecution}} updates its histograms 
> with {{toSeconds(startTime)}}, which is 
> {{TimeUnit.MILLISECONDS.toSeconds(...)}}. So the elapsed time is cut down to 
> whole seconds, and a reconcile under one second records 0.
> - Most reconciles take well under a second. So the quantiles of these 
> histograms are mostly 0, and their Prometheus {{_sum}} (SPARK-59935) 
> undercounts.
> - Recording nanoseconds in a histogram whose name contains {{nanos}} would 
> fix it, since {{PrometheusPullModelHandler}} then exports seconds. The name 
> change is user-facing, so it needs a migration guide entry.
> - This is there since SPARK-48984 (0.1.0).
> Found while reviewing 
> https://github.com/apache/spark-kubernetes-operator/pull/922.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to