Every service mesh benchmark you will find quotes a latency figure in the region of a fraction of a millisecond per hop, and every one of them is telling the truth. I measured Linkerd at +0.16 ms on a single connection, which sits comfortably inside the range its own documentation claims.
That number is also nearly useless, because it is measured at a concurrency nobody runs. At thirty-two concurrent connections the same setup lost 86% of its throughput.
Both figures come from the same cluster, the same afternoon, the same manifests with one annotation changed.
The setup
Three s-4vcpu-8gb nodes on DigitalOcean Kubernetes 1.36.3, and a deliberately boring application: nginx returning a fixed string at the backend, nginx reverse-proxying to it at the frontend. No database, no serialisation, no application logic at all.
That choice matters and I will come back to it, because it is the strongest objection to everything below.
The load generator is fortio, running as a pod inside the cluster so nothing in these numbers is internet latency. It is also what the Istio team use for their own benchmarks, which felt like the fair instrument to pick.
Two routes are measured. One hop goes straight from the load generator to the backend. Two hops goes through the frontend, which proxies onward. Comparing the two isolates what a single additional meshed hop costs rather than what an end-to-end request costs.
Then the whole thing runs twice, once bare and once with linkerd.io/inject=enabled on the namespace, on the same nodes, minutes apart.
What came out
| no mesh | Linkerd | ||
|---|---|---|---|
| one hop, c=1, p50 | 0.537 ms | 0.693 ms | +0.16 ms |
| one hop, c=1, throughput | 2,737 rps | 1,356 rps | -50% |
| one hop, c=32, p50 | 0.585 ms | 3.808 ms | +3.2 ms |
| one hop, c=32, throughput | 51,791 rps | 7,018 rps | -86% |
| two hops, c=32, p50 | 1.486 ms | 7.160 ms | +5.7 ms |
| two hops, c=32, throughput | 17,481 rps | 3,949 rps | -77% |
The latency columns tell a reassuring story and the throughput columns tell a different one. At a single connection the mesh adds a sixth of a millisecond, which is the figure that ends up on slides. At thirty-two it adds three and a quarter milliseconds to the median and takes six sevenths of the capacity.
Note what happens to the unmeshed baseline as concurrency rises: p50 stays flat at roughly half a millisecond from c=1 to c=32 while throughput climbs from 2,737 to 51,791 requests per second. That is what a system with headroom looks like. The meshed runs do the opposite, with latency climbing steadily as throughput refuses to.
The obvious objection, which I had too
My first instinct was that the proxies had simply run out of CPU, and that I had built a benchmark of my node size rather than of Linkerd.
Node utilisation during the meshed runs at c=32 peaked at 50%, with four cores per node:
nodes: 42% 48% 17%
Half the cluster idle, and throughput capped anyway. The ceiling is inside the proxy rather than on the machine.
I also checked whether Linkerd was throttling itself. There is no CPU limit set on the injected container and no core restriction in its environment, so it is taking what it asks for and the answer is not a configuration mistake I made.
To be thorough I ran the whole thing twice on different hardware. The first cluster had s-2vcpu-4gb nodes, where meshed throughput at c=32 was 4,888 requests per second. Doubling to four cores took it to 7,018, a real improvement. But the unmeshed baseline went from 33,779 to 51,791 over the same change, so the ratio barely moved. Bigger nodes buy you headroom, not a better exchange rate.
Where the CPU goes
This is the part I did not expect, and it explains the throughput ceiling better than the latency numbers do.
Under sustained load at c=32:
| pod | application | linkerd-proxy |
|---|---|---|
| backend | nginx 91-125m | 340-435m |
| frontend | nginx 273-395m | 627-688m |
| loadgen | fortio 218-264m | 559-596m |
The sidecar consistently used between two and three times the CPU of the container it was proxying for. Not a fraction, a multiple. On the backend pods, where nginx does almost nothing except return a constant, the proxy used roughly three and a third times as much CPU as the thing it exists to protect.
Memory is the opposite story, and it is genuinely impressive: 12 Mi per proxy, flat, whether idle or saturated. If you are sizing a cluster for memory, a mesh is close to free.
It does what it says
I want to be careful here, because the numbers above read as a prosecution and that is not what I found.
Every connection in the cluster came up mutually authenticated, with no application changes, no certificate management, and no configuration beyond a namespace annotation:
SRC DST SRC_NS DST_NS SECURED
frontend backend default default √
loadgen frontend default default √
Getting mTLS between every pair of services, with automatic certificate rotation, by annotating a namespace, is a genuinely remarkable piece of engineering. The question is not whether that is worth something. It obviously is. The question is whether the people choosing it know they are paying for it in throughput rather than in the latency figure they were shown.
What I got wrong
Four things, and the last one would have produced a conclusion in the opposite direction to everything above.
The backend would not start at all for the first few minutes. My nginx config had a return 200 '$BIGBODY' in it, left over from a payload-size test I abandoned, and nginx treats $BIGBODY as an undefined variable rather than a string, so the config failed validation and the container crash-looped. The load generator failed separately and for a dumber reason: the fortio image is distroless and has no sleep binary, so my command: ["sleep","infinity"] produced executable file not found in $PATH five times before I read the message properly.
Then kubectl top returned Metrics API not available, because DOKS does not ship metrics-server by default. I had already written down that the nodes were not saturated at that point, on no evidence whatsoever, and had to go back and install metrics-server before that sentence was worth anything.
The one that nearly cost me the finding came last. My first attempt at measuring sidecar CPU under load sampled kubectl top twelve seconds after starting the load, and metrics-server aggregates on a lag, so every proxy reported 1m of CPU. One milli-core. Had I stopped there I would have published that the sidecars are essentially free, which is the opposite of true, and I would have had a clean table to prove it. Sampling forty-five seconds into a seventy-five second run gives the 340 to 688m figures above, stable across three consecutive samples.
The strongest argument against my own numbers
nginx returning a fixed string is close to the worst possible case for a mesh, and I chose it deliberately to make the proxy's cost visible. When the application does nothing, the proxy's work is nearly all the work, so the relative tax is as large as it can possibly be.
A service that spends 50 ms waiting on a database will show something completely different. Add 3 ms of proxy to 50 ms of query and you have a 6% latency increase that nobody will ever notice, and the throughput ceiling moves too, because the application's own concurrency limits bind long before the proxy's do.
So do not read -86% as what your service will experience. Read it as the ceiling of what the proxy itself can push, which is the thing you are actually buying into, and then work out how close your service runs to that ceiling.
The services where this matters are the ones that look like my benchmark: internal RPC hops that do very little and get called constantly. Caches, auth checks, feature flag lookups, the small fast things sitting in the middle of every request path. Those are also, inconveniently, the services most likely to be meshed, because they are the ones talking to everything else.
What I would take from it
If you are evaluating a mesh, measure it at the concurrency you actually run, not at one connection, and measure throughput rather than only latency. The single-connection latency number is honest, reproducible, and the least informative thing in the whole exercise.
Budget CPU rather than memory. Twelve megabytes per pod is nothing. Two to three times your application's CPU, per pod, is a real line item and it is the one I have never seen on a slide.
And the mTLS is worth paying for in a great many environments. I would just rather people paid for it knowingly, having seen the number next to the one they were quoted.
The whole experiment, both clusters and all four benchmark runs, cost about 40 cents of cluster time and the manifests and raw fortio output are in the repository.

Top comments (2)
Useful numbers. The pattern I keep seeing is teams adopting a mesh for mTLS theater, then discovering the real cost at connection churn, not at idle. If your workload is short-lived and chatty, measure p99 under fan-out before you standardize on the mesh. Curious what workload shape produced the 86% hit in your test. iin1004h2128
One thing in the setup changes how I read the -86%. fortio sits in the meshed namespace, so the "one hop" route crosses two proxies, the loadgen's outbound and the backend's inbound. All 32 connections share that one outbound proxy, and none of the proxies in your table went past about 0.7 of a core while the nodes sat at half. That looks less like a CPU tax and more like a per-proxy ceiling. I'd rerun c=32 as four loadgen pods at c=8 each. If aggregate throughput climbs, the number to plan around is per pod rather than per cluster, and that is a different capacity conversation.