Live demo

The number you sized for is a six-week event.

90 days of hourly peak latency for five services, 2160 samples each. Everything below is a peaks-over-threshold fit computed in your browser: the shape parameter, the return levels, and the replica count that falls out of them.

checkout-apiPaymentsξ = 0.41 · heavy tail432 exceedances over 244ms
p99.9 of the sample
1.02s
what capacity is sized off
Worst hour in 90 days
2.06s
the whole sample's information about the tail
Modelled 1 in 365 days
2.94s
2.9× the p99.9
Modelled 1 in 1,000 days
4.41s
above anything ever observed
1.55s3.10s4.65s6.19sworst hour in 90 daysp99.9 · what you provisioned for1.12s30 d1.70s90 d2.94s365 d4.41s1,000 d1 day
What the shape parameter says

Heavy tail. Events far worse than anything observed are not just possible, they are expected on a long enough horizon. Size for the return level, not the maximum.

A p99.9 taken over hourly samples is one hour in a thousand. That is an event every 42 days by definition, before any modelling. The fitted tail puts it at 23 d. Either way it is not a rare event, and it is the number the capacity plan is built on.

What it costs in machines

A shape parameter you have never computed is setting your replica count

Little’s Law: in-flight requests are arrival rate times time in system. At 3,200 requests per second and 120 concurrent slots per replica, latency converts directly into machines. Nothing here is a fudge factor.

Sized for the p99.9
28 replicas
1.02s · 3,254 in flight
Sized for a yearly hour
79 replicas
2.94s · 9,422 in flight
Sized for 1 in 1,000 days
118 replicas
4.41s · 14,120 in flight

The gap between 28 and 79 is not a safety-margin argument. It is the difference between provisioning for the worst hour you happened to sample and provisioning for the worst hour a year contains. Sizing at 28 does not mean the yearly hour will not happen; it means requests get shed when it does, and the post-incident review calls it unprecedented traffic.

ξ is doing the work

Same percentile, opposite futures

Five services, each with a p99.9 their owners quote in planning meetings. The percentile tells you nothing about which of them has a ceiling. The shape parameter does, and it splits them cleanly: two extrapolate hard, one barely moves, and one is already at its ceiling.

Serviceξp99.9Worst hour1 in 365 dUnderstated byVerdict
checkout-api
0.41
1.02s2.06s2.94s2.90×Heavy. Size for the return level.
ledger-write
0.33
707ms1.04s1.39s1.96×Heavy. Size for the return level.
search-suggest
0.17
471ms542ms720ms1.53×Mild extrapolation, still real.
auth-token
0.01
185ms202ms225ms1.21×Exponential. The percentile is nearly right.
asset-edge
-0.26
115ms128ms122ms1.06×Bounded. Already at the ceiling.

Read the p99.9 column alone and asset-edge and checkout-api look like the same kind of problem at different scales. They are not. A negative ξ means a finite upper bound, and asset-edge has effectively already reached its. A positive ξ means there is no upper bound at all, and no amount of further observation will produce one.

The planning question

How bad does it get, and how many machines is that?

Pick a service and a horizon you are willing to be wrong on. The answer is a latency and a replica count, both derived from the fitted tail rather than from the worst thing you happened to record.

Fitted ξ
0.415
β = 52.9
Worst hour in 365 d
2.94s
p99.9 says 1.02s
Replicas required
79
28 if you size on the p99.9
Your p99.9 is exceeded
every 23 d
modelled recurrence of the level you sized for
To survive the worst hour in 365 days, checkout-api needs 79 replicas against the 28 a p99.9-based plan would buy. The shape parameter, not the percentile, is what moved that number.

Sized off the p99.9 of the last quarter. Pages roughly monthly and nobody knows why.