The problem
A service sits at 70% CPU utilization for months. p99 latency is a boring, flat 40ms. Someone decides 70% feels too close to the edge, "let's alert at 80% so we have room to react." Three weeks later the alert fires almost every afternoon during normal traffic peaks, pages people at 2am for nothing, and everyone starts snoozing it. Then one day utilization crosses 92% during a traffic bump that's maybe 15% above normal — not a spike, not an outage, just a slightly busier Tuesday — and p99 latency goes from 40ms to 800ms in about three minutes. Nobody changed anything. The system didn't get "a bit worse." It fell off a cliff.
This isn't a story about a bad on-call culture or a noisy alert. It's a story about treating utilization as a linear dial when it behaves nothing like one.
Why it happens
Model any single resource — a CPU core, a connection pool, a database — as a queue with one server and random arrivals (the standard M/M/1 queueing model). The expected time a request spends waiting in that queue, relative to how long it takes to actually get served, is:
wait multiplier = ρ / (1 − ρ)
where ρ (rho) is utilization, between 0 and 1.
That formula is a hyperbola, not a line. Here's what it actually produces:
| Utilization (ρ) | Wait multiplier |
|---|---|
| 50% | 1.0x |
| 60% | 1.5x |
| 70% | 2.33x |
| 80% | 4.0x |
| 90% | 9.0x |
| 95% | 19.0x |
| 99% | 99.0x |
Going from 70% to 90% utilization — a 20-point jump that sounds like "a bit more load" — multiplies queueing delay by roughly 4x (2.33x → 9x). Going from 90% to 95% does almost as much damage again. The denominator (1 − ρ) is heading to zero, and dividing by a number approaching zero is the entire story.
This is also why the damage looks instantaneous. Request arrivals aren't a smooth, even stream — they're bursty, roughly Poisson-distributed. Near saturation, a small burst has nowhere to drain to before the next one lands, so the queue compounds instead of recovering. Below ~70% utilization there's enough slack to absorb bursts between requests. Above ~90%, there isn't, so the same burst that used to be invisible now stacks on top of an already-full queue.
CPU percentage is just a visible proxy for ρ. The real quantity that matters is headroom: 1 − ρ. A dashboard showing "92% CPU" is really telling you your headroom shrank from 30% to 8% — a number that explains the multiplier a lot better than "92%" does on its own.
What to do about it
First, stop alerting on a flat utilization percentage. It mixes up two very different signals: "I'm busy" and "I'm about to fall off the curve." Alert on the thing you actually care about instead — p99 latency, queue depth, or request wait time directly — and treat utilization as a leading indicator you watch on a dashboard, not a page trigger.
Second, if you do keep a utilization-based alert (plenty of teams reasonably want an early-warning signal before latency visibly degrades), pick the threshold from the curve, not from a feeling. For latency-sensitive paths, the curve is already steep by 70% and brutal by 85%. "80% feels conservative" is backwards — 80% is already in the zone where a 10-point move in either direction swings delay by 2-3x.
Third, size your autoscaling buffer off the same math. If your scale-out trigger is slower than the time it takes ρ to cross from 80% to 95% under real traffic variance, the new capacity will land after the multiplier has already done its damage, not before.
A ten-line simulation makes the curve concrete:
def wait_multiplier(rho):
return rho / (1 - rho)
for rho in [0.5, 0.6, 0.7, 0.8, 0.9, 0.95, 0.99]:
print(f"{int(rho*100)}% -> {wait_multiplier(rho):.2f}x")
50% -> 1.00x
60% -> 1.50x
70% -> 2.33x
80% -> 4.00x
90% -> 9.00x
95% -> 19.00x
99% -> 99.00x
Run it once, and "80% feels safe" stops sounding reasonable.
Key takeaways
- Queueing delay scales as ρ/(1−ρ), a hyperbola — not linearly with utilization.
- Moving from 70% to 90% utilization isn't "20% more load," it's roughly 4x more queueing delay.
- The number that predicts pain is headroom (1−ρ), not the utilization percentage itself.
- Alert on latency and queue depth directly where you can; if you alert on utilization, pick the threshold from the curve, not a gut feeling.
- Autoscaling has to react faster than ρ crosses the steep part of the curve, or the new capacity arrives after the damage is done.
Top comments (0)