Hi Mukul,

Thanks for the positive feedback! It’s great to hear that you will consider 
adding higher percentiles in the next revision.
I completely agree with your point regarding P95 vs. P99 and the need to 
balance computation overhead with operational value. Focusing on P95 strikes 
the right balance—it effectively captures tail latencies and anomalous spikes 
without being overly sensitive to extreme 1% outliers.

Additionally, after reviewing the draft further, I'd like to share a few 
complementary observations regarding the current design for the WG and authors 
to consider:
1. Missing Time Window / Sampling Interval Context
Currently, the draft defines metrics like Max, Min, and Median (P50), but it 
does not specify the time window over which these statistics were calculated 
(e.g., is P50 computed over the last 5 seconds, 15 minutes, or since session 
reset?). Furthermore, sampling intervals are undefined.

  *   Risk: If the export frequency varies or is dynamically triggered, the BMP 
Collector/Server cannot accurately interpret these values. A "Max" over 5 
seconds has a drastically different operational meaning for capacity planning 
than a "Max" over 1 hour. Adding an optional Measurement Window Duration / 
Sampling Interval TLV would resolve this ambiguity.
2. Timestamp Precision Alignment (Seconds vs. Sub-seconds)
The current draft uses a 4-byte timestamp (second precision). Modern routing 
anomalies (e.g., BGP hijacking, rapid route leaks, or micro-bursts) and 
CPU/Queue spikes often occur at the millisecond or microsecond scale.

  *   Risk: 4-byte timestamps create a discrepancy with the standard 8-byte 
sub-second precision used in RFC 7854. More importantly, it introduces 
potential causality inversion—collectors may be unable to determine whether a 
CPU load spike preceded or followed a sudden burst of route updates. Aligning 
with sub-second timestamps would significantly enhance event correlation.
3. Lack of Standardization for Percentile Estimation Algorithms
The computation methodology for metrics like Median (P50) and P95 is currently 
a "black box." Different vendors (e.g., Cisco, Juniper, HPE, Huawei) may 
implement different estimation techniques (e.g., streaming algorithms like 
T-Digest, fixed-window sampling, or simplified binning/histogram methods).

  *   Risk: Heterogeneous calculation baselines across multi-vendor networks 
make cross-vendor telemetry aggregation and comparison unreliable for 
operators. It might be beneficial for the draft to either specify a recommended 
default algorithm or allow the router to signal the estimation method used in 
the metadata.
Regarding the router overhead you mentioned, these concerns could be addressed 
in the Operational Considerations section by making high percentiles 
optional/configurable, and recommending lightweight streaming estimation 
techniques (e.g., T-Digest or HdrHistogram) that require minimal memory and CPU 
footprints.

Looking forward to hearing your and the WG's thoughts on these points!

Best regards,
Shunwan


From: Srivastava, Mukul <[email protected]>
Sent: Thursday, August 13, 2026 10:09 PM
To: Zhuangshunwan <[email protected]>
Cc: [email protected]; [email protected]
Subject: Re: Mail regarding draft-ietf-grow-bmp-stats-informational-tlv

Hi Shunwan

Thanks for the review and comment.

We will explore this and update the draft to add higher percentiles in next 
version.
I am more inclined to keep P95 than P99 (which is almost 100%).

Note that maintaining/exporting these multiple percentiles stats might be an 
overhead at router. So we should maintain a balance between overhead and the 
value that it provides.

Thanks
Mukul

From: Zhuangshunwan 
<[email protected]<mailto:[email protected]>>
Date: Tuesday, August 11, 2026 at 2:27 AM
To: [email protected]<mailto:[email protected]> <[email protected]<mailto:[email protected]>>
Subject: [GROW] Mail regarding draft-ietf-grow-bmp-stats-informational-tlv
Hi Authors and GROW WG,

Thanks for introducing draft-ietf-grow-bmp-stats-informational-tlv-00. Moving 
to a flexible TLV structure for BMP Statistics is a very welcomed step.

I noticed that the draft currently introduces Median (P50) for statistical 
metrics. I would like to suggest exploring the addition of higher percentiles, 
such as P75, P95, or P99.

In operational telemetry (e.g., routing convergence time, processing latency, 
or queueing delays), metrics often exhibit long-tailed distributions. While P50 
provides a useful baseline, tail latencies and transient anomalies are usually 
masked by P50. In practice, operators rely heavily on P95 or P99 to detect SLA 
degradation and bottlenecks.

It would be great to get the authors' and the WG's thoughts on whether 
expanding the scope to cover higher percentiles makes sense for this draft.

Best regards,
Shunwan
_______________________________________________
GROW mailing list -- [email protected]
To unsubscribe send an email to [email protected]

Reply via email to