DEV Community

疏影
疏影

Posted on

Karpenter Consolidation Is Not Magic: The 4 Configs That Matter

Karpenter Consolidation Isn't Magic: The 4 Configs That Matter

We adopted Karpenter in March. Six months in, our cluster consolidation strategy lives in 4 configs and a Slack thread. Here's what actually moves the needle.

Why Karpenter over Cluster Autoscaler

We had 47 nodes running 24/7 with Cluster Autoscaler. CPU utilization was 18%. Cluster Autoscaler takes 4-8 minutes to provision a new node, and its consolidation logic is "wait 10 minutes, scale down if underutilized." That delay cost us 2-3 cold-start minutes per CI job during peak load.

Karpenter provisions in 45 seconds on average. It watches pod scheduling in real time and provisions the cheapest instance type that fits. Our ASG went from 47 nodes to 19 nodes, with 28% average CPU. Monthly bill dropped $3,200.

The 4 configs

1. EC2NodeClass — instance selection

apiVersion: karpenter.k8s.aws/v1beta1
kind: EC2NodeClass
spec:
  amiFamily: Bottlerocket
  blockDeviceMappings:
    - deviceName: /dev/xvda
        ebs:
          volumeSize: 100Gi
          volumeType: gp3
          iops: 3000
          encrypted: true
  securityGroupSelectorTerms:
    - tags:
        karpenter.sh/discovery: prod
  subnetSelectorTerms:
    - tags:
        karpenter.sh/discovery: prod
Enter fullscreen mode Exit fullscreen mode

Bottlerocket + gp3 is the cheap-and-fast combo. Spot instances dropped our compute cost 64% in steady state.

2. NodePool — what to run

apiVersion: karpenter.sh/v1beta1
kind: NodePool
spec:
  template:
    spec:
      requirements:
        - key: karpenter.sh/capacity-type
          operator: In
          values: ["spot", "on-demand"]
        - key: karpenter.k8s.aws/instance-category
          operator: In
          values: ["c", "m", "r"]
        - key: karpenter.k8s.aws/instance-generation
          operator: Gt
          values: ["4"]
  limits:
    cpu: "200"
    memory: 800Gi
  disruption:
    consolidationPolicy: WhenUnderutilized
    expireAfter: 720h
Enter fullscreen mode Exit fullscreen mode

Mixing spot with on-demand in the same NodePool means Karpenter picks spot first and falls back to on-demand when spot is unavailable. The expireAfter: 720h (30 days) forces node recycling for kernel updates.

3. Pod disruption budget

Critical: Karpenter's consolidation can be aggressive. Without a PDB, it can drain a deployment with 3 replicas simultaneously.

apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: api-pdb
spec:
  minAvailable: 2
  selector:
    matchLabels:
      app: api
Enter fullscreen mode Exit fullscreen mode

We learned this the hard way when Karpenter consolidated a 3-replica Redis cluster to a single node during a deploy. 14 minutes of 503s.

4. Consolidation TTL

disruption:
  consolidationPolicy: WhenUnderutilized
  expireAfter: 720h
Enter fullscreen mode Exit fullscreen mode

WhenUnderutilized is the default. WhenEmptyOrUnderutilized is safer for production — it never consolidates a node unless the node is empty OR every pod on it has another schedulable copy. We run WhenEmptyOrUnderutilized for prod, WhenUnderutilized for staging.

What still hurts

EBS volume thrash. Every consolidation cycle detaches and reattaches EBS volumes. For pods with 50GB stateful sets, this adds 90 seconds of pod startup. We now tag StatefulSets with karpenter.sh/do-not-disrupt: "true" and accept the cost.

Spot reclamation events. AWS gives you 2 minutes when reclaiming a Spot instance. Karpenter handles this via node-level eviction but the drain still races with your readiness probe. We added terminationGracePeriodSeconds: 90 to all production Deployments.

Multi-AZ bias. Karpenter will consolidate to a single AZ if that's the cheapest option. Bad for HA. We added topology.kubernetes.io/zone requirements.

The cost math

Our monthly AWS bill:

  • Before Karpenter: $12,400
  • After Karpenter: $7,100
  • Karpenter control plane cost: $0 (open source)

Net savings: $5,300/month. Payback was 1 week of engineering time.

On related tooling

If you're running Karpenter in a region with strict data residency (e.g., regulated finance), you'll want your CI agents to mount encrypted S3 buckets as local disks for cache. ScsDriver WebDAV mount tool for Windows handles this for Windows build agents — useful for hybrid Windows/Linux clusters where Windows nodes need to share build cache with Linux nodes. Pair it with a Karpenter NodePool that provisions r5.large instances for the Windows side.


What Karpenter configs are you tuning? Comment with your NodePool yaml.

Top comments (0)