DEV Community

Cover image for p99, Load Balancers and Autoscaling: Latency Intuition You Can Play With
DevOps Daily
DevOps Daily

Posted on

p99, Load Balancers and Autoscaling: Latency Intuition You Can Play With

Averages hide latency problems. Every SRE learns this, usually the hard way: the dashboard says 80ms average while support tickets say the app is slow, because the average smooths over a tail where one request in a hundred takes four seconds, and your heaviest users, the ones making the most requests, hit that tail most often.

Percentiles are how you see the tail, and the vocabulary around them (p50, p95, p99, tail latency) is quick to memorize and slow to internalize. Intuition, the ability to look at a p99 spike and have sensible suspects, usually comes from incidents. Three free browser simulators let you build a chunk of it without the incidents. Disclosure: I help build them; free, browser-based, no signup.

Part 1: watch a percentile move

The Latency Percentiles Simulator generates request-latency distributions and renders them as a histogram with P50, P90, P95 and P99 markers that move as you change the controls. The five scenarios are synthetic distributions shaped like five kinds of trouble:

  1. Healthy API: the baseline. A tight distribution; the percentiles sit close together. This is what "fine" looks like, and knowing its shape is what lets you recognize not-fine.
  2. Cold Starts: a small fraction of requests pay a big startup cost. P50 barely notices; P99 jumps. This is the canonical "the average is fine, the tail is not" case, and it is the shape serverless platforms and JIT warmup produce.
  3. Cache Churn: hits are fast, misses are slow, and the mix moves the middle percentiles around. Different shape than cold starts: the distribution goes wide rather than growing a spike.
  4. Noisy Neighbor: shared infrastructure steals capacity intermittently, smearing the whole distribution outward. The tail grows, but unlike cold starts, so does everything above the median.
  5. Rolling Deploy: a distribution shaped like a fleet mid-deploy, two populations blended, which is the shape to recognize in real telemetry when half your instances are new.

These are stylized shapes, not simulations of caches or neighbors, and that is fine for the purpose: the transferable skill is that similar p99 numbers can sit on very different distributions. Percentiles alone do not diagnose a cause, but knowing the classic shapes gives you better suspects to check when a tail spike arrives.

The other lesson hiding here is why tails matter more than their percentage suggests: if a page fans out to dozens of roughly independent backend calls and waits for all of them, the chance that at least one hits the p99 grows fast, so the "1% case" ends up in far more than 1% of page loads. Google's SRE material makes the same point: a backend's p99 can effectively become the frontend's median.

Part 2: the balancer decides who eats the tail

The Load Balancer Simulator animates a three-server fleet behind a balancer you configure live: the algorithm (round robin, least connections, IP hash, random), the traffic rate up to a burst mode, a server crash probability (up to a brutal 30%), retries on or off, and, the good part, you can click a server to knock it offline mid-run.

Two experiments teach the most. First, crank the failure rate with retries on and watch what a crash actually costs: the request does not just fail, it goes around again, adding load precisely when the fleet is struggling. Retries are not free; the counter of crashed-then-retried requests makes that concrete. Second, knock a server offline under each algorithm and watch the traffic redistribute; then try it with IP hash, where the same clients always land on the same server. Stickiness is what session-state architectures want, and the redistribution behavior when a sticky target dies is the price nobody mentions in the algorithm's one-line description.

Round robin versus least connections rounds it out: equal turns versus routing by current load. With identical healthy servers they look similar; under crashes and bursts, watching which algorithm piles work where is the point.

Part 3: capacity as a moving target

The Scaling Simulator hands you autoscaling controls (thresholds, instance counts, vertical versus horizontal choices) and four traffic scenarios to survive: Gradual Growth, Sudden Spike, Black Friday, Variable Load.

Gradual growth is the tutorial level; almost any policy works. The hard lessons are in the other three:

  • Sudden Spike punishes reactive scaling: by the time the threshold trips and instances come up, the spike already hurt users. The lesson is that autoscaling has a reaction time, and traffic that moves faster than it needs headroom, not thresholds.
  • Black Friday is the predictable-surge case: adding capacity before the surge you know is coming beats any reactive policy. The simulator lets you provision ahead by hand, which is the whole trick, minus the calendar automation real platforms add.
  • Variable Load is oscillating traffic, the pattern that punishes twitchy autoscaling with flapping in real systems. The simulator's scale-down cooldown is deliberately long, so what you actually observe is the defense working: capacity ratchets up and holds through the dips instead of thrashing. Cooldowns and stabilization windows are exactly how production autoscalers (Kubernetes' HPA included) buy calm at the cost of some idle capacity.

The through-line with parts 1 and 2: capacity decisions surface as latency and errors. The scaling simulator tracks response time and failures rather than full percentile distributions, but the connection stands: under-provisioned fleets are where the ugly histograms of part 1 come from, and the autoscaler's job is to buy the healthy shape with as few idle instances as possible.

Using the trio

A sequence that works: learn the five distribution shapes in part 1, run the crash-and-retry and kill-a-server experiments in part 2, then survive Sudden Spike and Black Friday in part 3. After that, three things in production read differently: a p99 chart makes you ask what the distribution under it looks like, balancer choice reads as a failure-behavior decision, and autoscaler settings read as a bet about how your traffic moves.

All three live with the rest of our free DevOps games and simulators. For the deep end afterwards, Google's SRE book chapter on monitoring distributed systems and Gil Tene's classic How NOT to Measure Latency are the standard references, easier to absorb once you have watched a histogram misbehave yourself.

Top comments (0)