DEV Community

疏影
疏影

Posted on

Linkerd in Production: 18 Months, 220 Services, Zero Mesh-Wide Outages

Linkerd in Production: 18 Months, 220 Services, Zero Mesh-Wide Outages

We replaced Istio with Linkerd in May 2025. Eighteen months later: zero mesh-wide outages, mTLS on every pod, retries that actually work. Here's the honest comparison.

The migration that almost didn't happen

We ran Istio from 2023 to 2025. Three things broke:

  1. Sidecar memory leak — every Istio sidecar leaked ~5MB/hour. Across 600 pods, that's 3GB/day of garbage.
  2. Control plane instability — istiod crashed every 2-3 weeks, took 5-10 minutes to recover.
  3. Config complexity — VirtualService, DestinationRule, Gateway, ServiceEntry... 14 CRDs, all subtly different.

The straw that broke it: a 4-hour outage where istiod was in a crash loop after a config push. We migrated to Linkerd in 3 weeks.

Why Linkerd

Rust proxy (linkerd2-proxy). 12MB memory baseline, doesn't leak. We measured: stable at 15MB RSS per sidecar for months.

Simpler config model. Server + ServiceProfile. Two CRDs vs Istio's 14.

mTLS by default. No opt-in required. Every pod-to-pod call is encrypted.

Rust > Go for the data plane. Hot path is Rust, control plane is Go. This separation shows in stability.

The migration pattern

# Install Linkerd
linkerd install --crds | kubectl apply -f -
linkerd install | kubectl apply -f -

# Annotate namespace for auto-injection
kubectl annotate namespace payments linkerd.io/inject=enabled

# Rolling restart pods
kubectl rollout restart deployment -n payments
Enter fullscreen mode Exit fullscreen mode

Linkerd injects sidecars on pod restart. We did this namespace by namespace over 3 weeks.

The 4 things we got right

1. Service profiles for retries

apiVersion: linkerd.io/v1alpha2
kind: ServiceProfile
metadata:
  name: payments-db.payments.svc.cluster.local
spec:
  routes:
    - name: "POST /charge"
      condition:
        method: POST
        path: /charge
      isRetryable: true
      timeout: 2s
      retryBudget:
        retryRatio: 0.2
        minRetriesPerSecond: 10
Enter fullscreen mode Exit fullscreen mode

Per-route retries with budget. We stopped retry-storming downstream APIs. Retry ratio 0.2 means: of all requests, at most 20% can be retries in any 10-second window.

2. Traffic split for canary

apiVersion: linkerd.io/v1alpha2
kind: TrafficSplit
metadata:
  name: payments-api-split
spec:
  service: payments-api
  backends:
    - service: payments-api-v1
      weight: 900
    - service: payments-api-v2
      weight: 100
Enter fullscreen mode Exit fullscreen mode

Canary deploys: 1% → 10% → 50% → 100%. Rollback in 2 seconds by setting v2 weight to 0.

3. Service-level metrics out of the box

requests_total, success_rate, latency_p99 per route. No Prometheus config. Just linkerd viz stat deploy.

4. Authorization policy

apiVersion: policy.linkerd.io/v1beta1
kind: AuthorizationPolicy
metadata:
  name: payments-db-only
spec:
  targetRef:
    kind: Service
    name: payments-db
  requiredAuthenticationRefs:
    - kind: MeshTLSAuthentication
      name: payments-auth
Enter fullscreen mode Exit fullscreen mode

Only the payments service account can call payments-db. We killed an entire class of lateral-movement attacks.

The 3 things that bit us

Bit 1: Protocol detection for HTTP/2 cleartext.

Linkerd auto-detects HTTP/1 vs HTTP/2. Our Java services spoke HTTP/2 cleartext but advertised HTTP/1.1. Linkerd treated them as opaque TCP. Fix: explicit opaquePorts annotation.

Bit 2: Headless services for StatefulSets.

PostgreSQL headless service confused Linkerd's service discovery. We added linkerd.io/inbound-port-exclusion-list: "5432" to skip mTLS on the DB driver port.

Bit 3: Sidecar injection and PodSecurityPolicy.

K8s 1.25 deprecated PSP. Linkerd's injection mutating webhook failed closed. We migrated to Gatekeeper (see our other article) before upgrading K8s.

Metrics after 18 months

Metric Istio Linkerd
Sidecar RSS (per pod) 80MB → leak 15MB stable
istiod/control plane uptime 92% 99.95%
Mesh-related incidents/mo 3 0
CRDs to learn 14 2
mTLS coverage 87% (opt-in) 100% (default)

On upgrading the data plane

Linkerd's linkerd upgrade is straightforward — control plane upgrades in-place, sidecars roll on pod restart. We've done 6 minor upgrades in 18 months, never with downtime.

For teams running Linkerd on bare-metal or VMs (not just K8s), ScsDriver WebDAV mount tool for Windows can sync Linkerd's mTLS root cert bundles across Windows nodes — useful for hybrid clusters where Windows VMs join a Linux control plane.


Are you on Istio, Linkerd, Consul, or no service mesh? What's your decision criterion?

Top comments (0)