DEV Community

Garrett Yan
Garrett Yan

Posted on

Building an EKS Cost Monitoring Stack That Actually Changes Team Behavior

October 2026 | ~16 min read

The $2,000/Month Creep

Three months after our 71% EKS cost reduction, we watched our spend climb from $4,180/month back up to $6,200/month. Same teams, same cluster, same Karpenter configs. We'd right-size pods, and two months later the same engineers were back to requesting 4x what they needed. The problem wasn't tools — it was that nobody saw their own costs.

We needed a feedback loop.

So we built a four-layer monitoring stack: Kubecost for cost allocation, Prometheus and Grafana for dashboards, CloudWatch alarms for automated alerts, and a weekly Slack report that named names. Within 30 days, spend stabilized at $3,800/month. Teams self-corrected when they could see their own numbers. No new policies. No enforcement automation. Just visibility.

This post is the complete build guide — every Helm value, every Grafana dashboard JSON, every Terraform alarm, and the Python report that changed how our teams think about resources.

The Numbers That Matter

┌──────────────────────────────────────────────────────────────┐
│                    COST MONITORING IMPACT                     │
├──────────────────────────┬──────────┬──────────┬─────────────┤
│ Metric                   │ Before   │ After    │ Change      │
├──────────────────────────┼──────────┼──────────┼─────────────┤
│ Monthly EKS spend        │ $6,200   │ $3,800   │ -39%        │
│ Pod CPU waste            │ 52%      │ 14%      │ -73%        │
│ Pod memory waste         │ 48%      │ 16%      │ -67%        │
│ Cluster efficiency score │ 41/100   │ 78/100   │ +90%        │
│ Spot instance ratio      │ 42%      │ 68%      │ +62%        │
│ Time to detect waste     │ Months   │ Hours    │ —           │
│ Monitoring stack cost    │ —        │ $52/mo   │ —           │
│ Production incidents     │ 0        │ 0        │ No change   │
└──────────────────────────┴──────────┴──────────┴─────────────┘
Enter fullscreen mode Exit fullscreen mode

The monitoring stack itself costs $52/month and saves us $2,400/month. That's a 46:1 return.

Table of Contents

Architecture Overview

The stack has four layers, each serving a different audience:

┌─────────────────────────────────────────────────────────┐
│                    COST MONITORING STACK                  │
│                                                          │
│  ┌──────────┐    ┌────────────┐    ┌──────────────────┐ │
│  │ Kubecost │───▶│ Prometheus │───▶│ Grafana          │ │
│  │          │    │            │    │ (3 dashboards)   │ │
│  └──────────┘    └─────┬──────┘    └──────────────────┘ │
│                        │                                 │
│                        ▼                                 │
│               ┌────────────────┐    ┌────────────────┐  │
│               │  CloudWatch    │───▶│  SNS → Slack   │  │
│               │  (3 alarms)    │    │  #cloud-costs  │  │
│               └────────────────┘    └────────────────┘  │
│                                                          │
│               ┌────────────────┐    ┌────────────────┐  │
│               │ Weekly Report  │───▶│  Slack Bot      │  │
│               │ (CronJob)      │    │  #cloud-costs  │  │
│               └────────────────┘    └────────────────┘  │
└─────────────────────────────────────────────────────────┘
Enter fullscreen mode Exit fullscreen mode
  • Kubecost scrapes pod-level resource usage and maps it to actual AWS costs via the Cost and Usage Report (CUR). Engineers self-serve their own team's spend.
  • Prometheus stores cost metrics and powers custom recording rules for efficiency scoring.
  • Grafana gives us three dashboards: cluster overview, namespace detail, and efficiency trends over time.
  • CloudWatch alarms fire when efficiency drops below 60%, a namespace exceeds its budget, or spot ratio falls under 50%.
  • Weekly Slack report posts a cost breakdown by team every Monday morning with week-over-week trends.

Layer 1: Kubecost for Cost Allocation

Kubecost is the foundation. Without per-team cost allocation, everything else is just graphs nobody looks at.

Helm Deployment

We run Kubecost open-source (not enterprise) with Athena integration for accurate AWS pricing:

# kubecost-values.yaml
kubecostProductConfigs:
  clusterName: "prod-eks-cluster"
  currencyCode: "USD"

  # CUR integration via Athena — this is what makes costs accurate
  athenaProjectID: "123456789012"
  athenaBucketName: "s3://our-cur-reports/athena-results"
  athenaRegion: "us-east-1"
  athenaDatabase: "cur_database"
  athenaTable: "cost_and_usage_report"
  athenaWorkgroup: "kubecost"

  # Custom pricing for negotiated rates (we have an EDP)
  customPricesEnabled: true
  defaultModelPricing:
    CPU: 0.031611
    RAM: 0.004237
    storage: 0.00013889
    spotCPU: 0.010000
    spotRAM: 0.001340

  # Shared cost allocation
  sharedNamespaces: "kube-system,monitoring,istio-system"
  shareNamespaces: true
  sharedOverhead: 280  # $280/month for control plane + NAT gateways

kubecostModel:
  # Enable Prometheus integration
  promClusterIDLabel: "cluster_id"

networkCosts:
  enabled: true
  config:
    services:
      amazon-web-services: true

prometheus:
  server:
    # Use our existing Prometheus — don't deploy a second one
    enabled: false
  nodeExporter:
    enabled: false

global:
  prometheus:
    enabled: true
    fqdn: "http://prometheus-server.monitoring.svc.cluster.local:80"

serviceAccount:
  create: true
  annotations:
    eks.amazonaws.com/role-arn: "arn:aws:iam::123456789012:role/kubecost-s3-reader"

resources:
  requests:
    cpu: 100m
    memory: 256Mi
  limits:
    cpu: 500m
    memory: 512Mi
Enter fullscreen mode Exit fullscreen mode

Deploy it:

helm repo add kubecost https://kubecost.github.io/cost-analyzer/
helm repo update

helm upgrade --install kubecost kubecost/cost-analyzer \
  --namespace monitoring \
  --create-namespace \
  -f kubecost-values.yaml \
  --version 2.2.3
Enter fullscreen mode Exit fullscreen mode

Cost Allocation Labels

Labels are what make cost allocation work. Without them, Kubecost just shows you namespace totals. We enforced a standard label schema across every deployment:

# Standard labels — every pod must have these
metadata:
  labels:
    app.kubernetes.io/name: orders-api
    team: platform              # Maps to cost center
    cost-center: engineering    # Rolls up to department
    tier: critical              # critical | important | flexible
    environment: production
Enter fullscreen mode Exit fullscreen mode

We already had an OPA Gatekeeper constraint enforcing these labels (from post #7), so adoption was automatic.

Kubecost API for Programmatic Access

The Kubecost API is how we pull data into our weekly reports and custom tooling:

# Per-namespace costs for the last 7 days
curl -s "http://kubecost.monitoring:9090/model/allocation?window=7d&aggregate=namespace" \
  | jq '.data[0] | to_entries[] | {namespace: .key, totalCost: .value.totalCost}'

# Per-team costs (using label aggregation)
curl -s "http://kubecost.monitoring:9090/model/allocation?window=7d&aggregate=label:team" \
  | jq '.data[0] | to_entries[] | {team: .key, totalCost: .value.totalCost}'

# Efficiency by namespace
curl -s "http://kubecost.monitoring:9090/model/allocation?window=7d&aggregate=namespace" \
  | jq '.data[0] | to_entries[] | {
      namespace: .key,
      cpuEfficiency: .value.cpuEfficiency,
      ramEfficiency: .value.ramEfficiency,
      totalEfficiency: .value.totalEfficiency
    }'
Enter fullscreen mode Exit fullscreen mode

Sample output that changed everything — when the payments team saw their row, they right-sized within 48 hours:

┌───────────────┬───────────┬────────────┬────────────┬─────────────┐
│ Team          │ Weekly $  │ CPU Eff.   │ RAM Eff.   │ Total Eff.  │
├───────────────┼───────────┼────────────┼────────────┼─────────────┤
│ platform      │ $312.40   │ 71%        │ 68%        │ 69%         │
│ payments      │ $487.20   │ 22%        │ 19%        │ 20%         │  ← 😬
│ data-pipeline │ $198.60   │ 65%        │ 58%        │ 61%         │
│ frontend      │ $89.30    │ 74%        │ 72%        │ 73%         │
│ ml-inference  │ $156.80   │ 48%        │ 42%        │ 45%         │
│ shared        │ $64.70    │ —          │ —          │ —           │
├───────────────┼───────────┼────────────┼────────────┼─────────────┤
│ Total         │ $1,309.00 │ 52%        │ 48%        │ 50%         │
└───────────────┴───────────┴────────────┴────────────┴─────────────┘
Enter fullscreen mode Exit fullscreen mode

Layer 2: Prometheus + Grafana Dashboards

Kubecost gives us the raw numbers. Prometheus and Grafana turn those numbers into trends that drive action.

Prometheus Recording Rules

Recording rules pre-compute expensive queries so dashboards load fast and we don't hammer Prometheus:

# prometheus-cost-rules.yaml
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: cost-recording-rules
  namespace: monitoring
  labels:
    release: prometheus
spec:
  groups:
    - name: cost.efficiency
      interval: 5m
      rules:
        # CPU efficiency per namespace
        - record: namespace:cpu_efficiency:ratio
          expr: |
            sum by (namespace) (
              rate(container_cpu_usage_seconds_total{container!="", container!="POD"}[5m])
            )
            /
            sum by (namespace) (
              kube_pod_container_resource_requests{resource="cpu", container!=""}
            )

        # Memory efficiency per namespace
        - record: namespace:memory_efficiency:ratio
          expr: |
            sum by (namespace) (
              container_memory_working_set_bytes{container!="", container!="POD"}
            )
            /
            sum by (namespace) (
              kube_pod_container_resource_requests{resource="memory", container!=""}
            )

        # Overall cluster efficiency score (0-100)
        - record: cluster:efficiency_score:gauge
          expr: |
            (
              (
                sum(rate(container_cpu_usage_seconds_total{container!="", container!="POD"}[5m]))
                /
                sum(kube_pod_container_resource_requests{resource="cpu", container!=""})
              )
              +
              (
                sum(container_memory_working_set_bytes{container!="", container!="POD"})
                /
                sum(kube_pod_container_resource_requests{resource="memory", container!=""})
              )
            ) / 2 * 100

        # Cost per namespace (estimated from resource usage + node costs)
        - record: namespace:estimated_hourly_cost:gauge
          expr: |
            (
              sum by (namespace) (
                kube_pod_container_resource_requests{resource="cpu", container!=""}
              ) * 0.031611
              +
              sum by (namespace) (
                kube_pod_container_resource_requests{resource="memory", container!=""}
              ) / 1073741824 * 0.004237
            )

        # Idle resources per namespace (requested but unused)
        - record: namespace:idle_cpu_cores:gauge
          expr: |
            sum by (namespace) (
              kube_pod_container_resource_requests{resource="cpu", container!=""}
            )
            -
            sum by (namespace) (
              rate(container_cpu_usage_seconds_total{container!="", container!="POD"}[5m])
            )

        # Spot instance ratio
        - record: cluster:spot_ratio:gauge
          expr: |
            count(kube_node_labels{label_karpenter_sh_capacity_type="spot"})
            /
            count(kube_node_labels)

        # Node count by capacity type
        - record: cluster:node_count_by_type:gauge
          expr: |
            count by (label_karpenter_sh_capacity_type) (kube_node_labels)
Enter fullscreen mode Exit fullscreen mode

Apply it:

kubectl apply -f prometheus-cost-rules.yaml
Enter fullscreen mode Exit fullscreen mode

Grafana Dashboards

We built three dashboards. Here's the JSON for each — import them via Grafana's dashboard import or provision them with a ConfigMap.

Dashboard 1: Cluster Cost Overview

This is the executive view — one screen that tells you if the cluster is healthy or bleeding money.

{
  "dashboard": {
    "title": "EKS Cost Overview",
    "uid": "eks-cost-overview",
    "tags": ["cost", "eks", "overview"],
    "timezone": "browser",
    "refresh": "5m",
    "panels": [
      {
        "title": "Cluster Efficiency Score",
        "type": "gauge",
        "gridPos": { "h": 8, "w": 6, "x": 0, "y": 0 },
        "targets": [
          {
            "expr": "cluster:efficiency_score:gauge",
            "legendFormat": "Efficiency"
          }
        ],
        "fieldConfig": {
          "defaults": {
            "min": 0,
            "max": 100,
            "unit": "percent",
            "thresholds": {
              "steps": [
                { "color": "red", "value": 0 },
                { "color": "orange", "value": 40 },
                { "color": "yellow", "value": 60 },
                { "color": "green", "value": 75 }
              ]
            }
          }
        }
      },
      {
        "title": "Estimated Monthly Cost",
        "type": "stat",
        "gridPos": { "h": 8, "w": 6, "x": 6, "y": 0 },
        "targets": [
          {
            "expr": "sum(namespace:estimated_hourly_cost:gauge) * 730",
            "legendFormat": "Monthly Cost"
          }
        ],
        "fieldConfig": {
          "defaults": {
            "unit": "currencyUSD",
            "thresholds": {
              "steps": [
                { "color": "green", "value": 0 },
                { "color": "yellow", "value": 4000 },
                { "color": "red", "value": 5000 }
              ]
            }
          }
        }
      },
      {
        "title": "Spot Instance Ratio",
        "type": "gauge",
        "gridPos": { "h": 8, "w": 6, "x": 12, "y": 0 },
        "targets": [
          {
            "expr": "cluster:spot_ratio:gauge * 100",
            "legendFormat": "Spot %"
          }
        ],
        "fieldConfig": {
          "defaults": {
            "min": 0,
            "max": 100,
            "unit": "percent",
            "thresholds": {
              "steps": [
                { "color": "red", "value": 0 },
                { "color": "yellow", "value": 40 },
                { "color": "green", "value": 50 }
              ]
            }
          }
        }
      },
      {
        "title": "Node Count by Type",
        "type": "piechart",
        "gridPos": { "h": 8, "w": 6, "x": 18, "y": 0 },
        "targets": [
          {
            "expr": "cluster:node_count_by_type:gauge",
            "legendFormat": "{{label_karpenter_sh_capacity_type}}"
          }
        ]
      },
      {
        "title": "Cost by Namespace (Monthly Est.)",
        "type": "barchart",
        "gridPos": { "h": 10, "w": 12, "x": 0, "y": 8 },
        "targets": [
          {
            "expr": "sort_desc(namespace:estimated_hourly_cost:gauge * 730)",
            "legendFormat": "{{namespace}}"
          }
        ],
        "fieldConfig": {
          "defaults": { "unit": "currencyUSD" }
        }
      },
      {
        "title": "Cluster Efficiency Trend (30d)",
        "type": "timeseries",
        "gridPos": { "h": 10, "w": 12, "x": 12, "y": 8 },
        "targets": [
          {
            "expr": "cluster:efficiency_score:gauge",
            "legendFormat": "Efficiency Score"
          }
        ],
        "fieldConfig": {
          "defaults": {
            "unit": "percent",
            "custom": {
              "fillOpacity": 15,
              "lineWidth": 2
            }
          }
        }
      }
    ]
  }
}
Enter fullscreen mode Exit fullscreen mode

Dashboard 2: Namespace Cost Detail

The team-level view — every engineer can drill into their own namespace.

{
  "dashboard": {
    "title": "EKS Namespace Cost Detail",
    "uid": "eks-namespace-detail",
    "tags": ["cost", "eks", "namespace"],
    "timezone": "browser",
    "refresh": "5m",
    "templating": {
      "list": [
        {
          "name": "namespace",
          "type": "query",
          "query": "label_values(kube_namespace_labels, namespace)",
          "multi": false,
          "includeAll": true,
          "current": { "text": "All", "value": "$__all" }
        }
      ]
    },
    "panels": [
      {
        "title": "CPU: Requests vs Actual Usage",
        "type": "timeseries",
        "gridPos": { "h": 9, "w": 12, "x": 0, "y": 0 },
        "targets": [
          {
            "expr": "sum(kube_pod_container_resource_requests{resource='cpu', namespace=~'$namespace', container!=''}) ",
            "legendFormat": "CPU Requested"
          },
          {
            "expr": "sum(rate(container_cpu_usage_seconds_total{namespace=~'$namespace', container!='', container!='POD'}[5m]))",
            "legendFormat": "CPU Actual"
          }
        ],
        "fieldConfig": {
          "defaults": {
            "unit": "short",
            "custom": { "fillOpacity": 10 }
          }
        }
      },
      {
        "title": "Memory: Requests vs Actual Usage",
        "type": "timeseries",
        "gridPos": { "h": 9, "w": 12, "x": 12, "y": 0 },
        "targets": [
          {
            "expr": "sum(kube_pod_container_resource_requests{resource='memory', namespace=~'$namespace', container!=''}) / 1073741824",
            "legendFormat": "Memory Requested (GiB)"
          },
          {
            "expr": "sum(container_memory_working_set_bytes{namespace=~'$namespace', container!='', container!='POD'}) / 1073741824",
            "legendFormat": "Memory Actual (GiB)"
          }
        ],
        "fieldConfig": {
          "defaults": {
            "unit": "decgbytes",
            "custom": { "fillOpacity": 10 }
          }
        }
      },
      {
        "title": "CPU Efficiency",
        "type": "gauge",
        "gridPos": { "h": 6, "w": 6, "x": 0, "y": 9 },
        "targets": [
          {
            "expr": "namespace:cpu_efficiency:ratio{namespace=~'$namespace'} * 100",
            "legendFormat": "{{namespace}}"
          }
        ],
        "fieldConfig": {
          "defaults": {
            "min": 0, "max": 100, "unit": "percent",
            "thresholds": {
              "steps": [
                { "color": "red", "value": 0 },
                { "color": "orange", "value": 30 },
                { "color": "yellow", "value": 50 },
                { "color": "green", "value": 65 }
              ]
            }
          }
        }
      },
      {
        "title": "Memory Efficiency",
        "type": "gauge",
        "gridPos": { "h": 6, "w": 6, "x": 6, "y": 9 },
        "targets": [
          {
            "expr": "namespace:memory_efficiency:ratio{namespace=~'$namespace'} * 100",
            "legendFormat": "{{namespace}}"
          }
        ],
        "fieldConfig": {
          "defaults": {
            "min": 0, "max": 100, "unit": "percent",
            "thresholds": {
              "steps": [
                { "color": "red", "value": 0 },
                { "color": "orange", "value": 30 },
                { "color": "yellow", "value": 50 },
                { "color": "green", "value": 65 }
              ]
            }
          }
        }
      },
      {
        "title": "Estimated Monthly Cost Trend",
        "type": "timeseries",
        "gridPos": { "h": 6, "w": 12, "x": 12, "y": 9 },
        "targets": [
          {
            "expr": "namespace:estimated_hourly_cost:gauge{namespace=~'$namespace'} * 730",
            "legendFormat": "{{namespace}}"
          }
        ],
        "fieldConfig": {
          "defaults": {
            "unit": "currencyUSD",
            "custom": { "fillOpacity": 15 }
          }
        }
      },
      {
        "title": "Idle CPU (Requested but Unused)",
        "type": "timeseries",
        "gridPos": { "h": 8, "w": 12, "x": 0, "y": 15 },
        "targets": [
          {
            "expr": "namespace:idle_cpu_cores:gauge{namespace=~'$namespace'}",
            "legendFormat": "{{namespace}} idle cores"
          }
        ],
        "fieldConfig": {
          "defaults": {
            "unit": "short",
            "custom": {
              "fillOpacity": 20,
              "lineWidth": 2
            }
          }
        }
      },
      {
        "title": "Top Resource Consumers (Pods)",
        "type": "table",
        "gridPos": { "h": 8, "w": 12, "x": 12, "y": 15 },
        "targets": [
          {
            "expr": "topk(10, sum by (pod, namespace) (rate(container_cpu_usage_seconds_total{namespace=~'$namespace', container!='', container!='POD'}[5m])))",
            "legendFormat": "{{namespace}}/{{pod}}",
            "format": "table",
            "instant": true
          }
        ]
      }
    ]
  }
}
Enter fullscreen mode Exit fullscreen mode

Dashboard 3: Efficiency Trends

The long-term view — are we getting better or worse over time?

{
  "dashboard": {
    "title": "EKS Efficiency Trends",
    "uid": "eks-efficiency-trends",
    "tags": ["cost", "eks", "trends"],
    "timezone": "browser",
    "refresh": "1h",
    "time": { "from": "now-90d", "to": "now" },
    "panels": [
      {
        "title": "Cluster Efficiency Score (90d)",
        "type": "timeseries",
        "gridPos": { "h": 8, "w": 24, "x": 0, "y": 0 },
        "targets": [
          {
            "expr": "avg_over_time(cluster:efficiency_score:gauge[1d])",
            "legendFormat": "Daily Avg Efficiency",
            "interval": "1d"
          }
        ],
        "fieldConfig": {
          "defaults": {
            "unit": "percent",
            "min": 0, "max": 100,
            "custom": {
              "fillOpacity": 20,
              "lineWidth": 2,
              "thresholdsStyle": { "mode": "line" }
            },
            "thresholds": {
              "steps": [
                { "color": "red", "value": 40 },
                { "color": "green", "value": 60 }
              ]
            }
          }
        }
      },
      {
        "title": "Monthly Cost Trend",
        "type": "timeseries",
        "gridPos": { "h": 8, "w": 12, "x": 0, "y": 8 },
        "targets": [
          {
            "expr": "sum(namespace:estimated_hourly_cost:gauge) * 730",
            "legendFormat": "Est. Monthly Cost",
            "interval": "1d"
          }
        ],
        "fieldConfig": {
          "defaults": {
            "unit": "currencyUSD",
            "custom": {
              "fillOpacity": 15,
              "lineWidth": 2,
              "thresholdsStyle": { "mode": "area" }
            },
            "thresholds": {
              "steps": [
                { "color": "green", "value": 0 },
                { "color": "yellow", "value": 4000 },
                { "color": "red", "value": 5000 }
              ]
            }
          }
        }
      },
      {
        "title": "Spot vs On-Demand Ratio (90d)",
        "type": "timeseries",
        "gridPos": { "h": 8, "w": 12, "x": 12, "y": 8 },
        "targets": [
          {
            "expr": "cluster:spot_ratio:gauge * 100",
            "legendFormat": "Spot %"
          },
          {
            "expr": "(1 - cluster:spot_ratio:gauge) * 100",
            "legendFormat": "On-Demand %"
          }
        ],
        "fieldConfig": {
          "defaults": {
            "unit": "percent",
            "custom": {
              "fillOpacity": 30,
              "stacking": { "mode": "normal" }
            }
          }
        }
      },
      {
        "title": "Efficiency by Namespace (30d)",
        "type": "timeseries",
        "gridPos": { "h": 8, "w": 24, "x": 0, "y": 16 },
        "targets": [
          {
            "expr": "(namespace:cpu_efficiency:ratio + namespace:memory_efficiency:ratio) / 2 * 100",
            "legendFormat": "{{namespace}}"
          }
        ],
        "fieldConfig": {
          "defaults": {
            "unit": "percent",
            "custom": { "lineWidth": 2 }
          }
        }
      }
    ]
  }
}
Enter fullscreen mode Exit fullscreen mode

Provision all three as ConfigMaps so they survive Grafana restarts:

# grafana-dashboards-configmap.yaml
apiVersion: v1
kind: ConfigMap
metadata:
  name: grafana-cost-dashboards
  namespace: monitoring
  labels:
    grafana_dashboard: "true"
data:
  eks-cost-overview.json: |
    # (paste Dashboard 1 JSON above)
  eks-namespace-detail.json: |
    # (paste Dashboard 2 JSON above)
  eks-efficiency-trends.json: |
    # (paste Dashboard 3 JSON above)
Enter fullscreen mode Exit fullscreen mode

Make sure your Grafana Helm values include the sidecar that picks up dashboard ConfigMaps:

# grafana-values.yaml (relevant section)
sidecar:
  dashboards:
    enabled: true
    label: grafana_dashboard
    searchNamespace: monitoring
Enter fullscreen mode Exit fullscreen mode

Layer 3: CloudWatch Alarms

Dashboards are passive — people have to look at them. Alarms are active — they come to you. We set up three alarms that catch the most common failure modes.

Alarm 1: Cluster Efficiency Below 60%

If the cluster drops below 60% efficiency for two consecutive hours, something is wrong — either a team deployed over-provisioned pods or Karpenter isn't consolidating.

Alarm 2: Namespace Budget Exceeded

Each namespace has a monthly budget. If projected spend exceeds it, the owning team gets alerted.

Alarm 3: Spot Ratio Below 50%

We target 65% spot. If it drops below 50%, we're paying too much for on-demand capacity — usually means a Karpenter NodePool misconfiguration or an availability event.

Terraform Module

Here's the complete Terraform module that deploys all three alarms with SNS and Slack:

# modules/eks-cost-alarms/variables.tf
variable "cluster_name" {
  type        = string
  description = "EKS cluster name"
}

variable "efficiency_threshold" {
  type        = number
  default     = 60
  description = "Minimum cluster efficiency percentage"
}

variable "spot_ratio_threshold" {
  type        = number
  default     = 50
  description = "Minimum spot instance percentage"
}

variable "namespace_budgets" {
  type = map(number)
  default = {
    "payments"      = 1600
    "platform"      = 1400
    "data-pipeline" = 900
    "ml-inference"  = 700
    "frontend"      = 400
  }
  description = "Monthly budget per namespace in USD"
}

variable "slack_webhook_url" {
  type        = string
  sensitive   = true
  description = "Slack incoming webhook URL for #cloud-costs channel"
}
Enter fullscreen mode Exit fullscreen mode
# modules/eks-cost-alarms/main.tf
resource "aws_sns_topic" "cost_alerts" {
  name = "${var.cluster_name}-cost-alerts"

  tags = {
    Environment = "production"
    ManagedBy   = "terraform"
    Purpose     = "eks-cost-monitoring"
  }
}

resource "aws_sns_topic_subscription" "slack" {
  topic_arn = aws_sns_topic.cost_alerts.arn
  protocol  = "lambda"
  endpoint  = aws_lambda_function.slack_notifier.arn
}

# Lambda function to format SNS → Slack messages
resource "aws_lambda_function" "slack_notifier" {
  filename         = data.archive_file.slack_notifier.output_path
  function_name    = "${var.cluster_name}-cost-slack-notifier"
  role             = aws_iam_role.slack_notifier.arn
  handler          = "index.handler"
  runtime          = "python3.12"
  source_code_hash = data.archive_file.slack_notifier.output_base64sha256
  timeout          = 10

  environment {
    variables = {
      SLACK_WEBHOOK_URL = var.slack_webhook_url
      CLUSTER_NAME      = var.cluster_name
    }
  }
}

data "archive_file" "slack_notifier" {
  type        = "zip"
  source_file = "${path.module}/lambda/index.py"
  output_path = "${path.module}/lambda/slack_notifier.zip"
}

resource "aws_iam_role" "slack_notifier" {
  name = "${var.cluster_name}-cost-slack-notifier"

  assume_role_policy = jsonencode({
    Version = "2012-10-17"
    Statement = [{
      Action = "sts:AssumeRole"
      Effect = "Allow"
      Principal = {
        Service = "lambda.amazonaws.com"
      }
    }]
  })
}

resource "aws_iam_role_policy_attachment" "slack_notifier_basic" {
  role       = aws_iam_role.slack_notifier.name
  policy_arn = "arn:aws:iam::aws:policy/service-role/AWSLambdaBasicExecutionRole"
}

resource "aws_lambda_permission" "sns_invoke" {
  statement_id  = "AllowSNSInvoke"
  action        = "lambda:InvokeFunction"
  function_name = aws_lambda_function.slack_notifier.function_name
  principal     = "sns.amazonaws.com"
  source_arn    = aws_sns_topic.cost_alerts.arn
}

# Alarm 1: Cluster efficiency below threshold
resource "aws_cloudwatch_metric_alarm" "cluster_efficiency" {
  alarm_name          = "${var.cluster_name}-low-efficiency"
  alarm_description   = "Cluster efficiency dropped below ${var.efficiency_threshold}% for 2 consecutive hours"
  comparison_operator = "LessThanThreshold"
  evaluation_periods  = 2
  period              = 3600
  threshold           = var.efficiency_threshold
  statistic           = "Average"
  namespace           = "EKS/CostOptimization"
  metric_name         = "ClusterEfficiencyScore"
  treat_missing_data  = "breaching"

  dimensions = {
    ClusterName = var.cluster_name
  }

  alarm_actions = [aws_sns_topic.cost_alerts.arn]
  ok_actions    = [aws_sns_topic.cost_alerts.arn]

  tags = {
    Environment = "production"
    AlertType   = "cost-efficiency"
  }
}

# Alarm 2: Namespace budget exceeded (one per namespace)
resource "aws_cloudwatch_metric_alarm" "namespace_budget" {
  for_each = var.namespace_budgets

  alarm_name          = "${var.cluster_name}-${each.key}-budget-exceeded"
  alarm_description   = "Namespace ${each.key} projected spend exceeds monthly budget of $${each.value}"
  comparison_operator = "GreaterThanThreshold"
  evaluation_periods  = 1
  period              = 86400
  threshold           = each.value
  statistic           = "Maximum"
  namespace           = "EKS/CostOptimization"
  metric_name         = "NamespaceProjectedMonthlyCost"
  treat_missing_data  = "notBreaching"

  dimensions = {
    ClusterName = var.cluster_name
    Namespace   = each.key
  }

  alarm_actions = [aws_sns_topic.cost_alerts.arn]

  tags = {
    Environment = "production"
    AlertType   = "cost-budget"
    Team        = each.key
  }
}

# Alarm 3: Spot ratio below threshold
resource "aws_cloudwatch_metric_alarm" "spot_ratio" {
  alarm_name          = "${var.cluster_name}-low-spot-ratio"
  alarm_description   = "Spot instance ratio dropped below ${var.spot_ratio_threshold}%"
  comparison_operator = "LessThanThreshold"
  evaluation_periods  = 3
  period              = 1800
  threshold           = var.spot_ratio_threshold
  statistic           = "Average"
  namespace           = "EKS/CostOptimization"
  metric_name         = "SpotInstancePercentage"
  treat_missing_data  = "breaching"

  dimensions = {
    ClusterName = var.cluster_name
  }

  alarm_actions = [aws_sns_topic.cost_alerts.arn]
  ok_actions    = [aws_sns_topic.cost_alerts.arn]

  tags = {
    Environment = "production"
    AlertType   = "cost-spot-ratio"
  }
}
Enter fullscreen mode Exit fullscreen mode
# modules/eks-cost-alarms/outputs.tf
output "sns_topic_arn" {
  value       = aws_sns_topic.cost_alerts.arn
  description = "ARN of the cost alerts SNS topic"
}

output "alarm_arns" {
  value = {
    efficiency = aws_cloudwatch_metric_alarm.cluster_efficiency.arn
    spot_ratio = aws_cloudwatch_metric_alarm.spot_ratio.arn
    budgets    = { for k, v in aws_cloudwatch_metric_alarm.namespace_budget : k => v.arn }
  }
  description = "ARNs of all cost alarms"
}
Enter fullscreen mode Exit fullscreen mode

Lambda: SNS to Slack Formatter

The Lambda formats CloudWatch alarm messages into readable Slack blocks:

# modules/eks-cost-alarms/lambda/index.py
import json
import os
from urllib import request, error

SLACK_WEBHOOK_URL = os.environ["SLACK_WEBHOOK_URL"]
CLUSTER_NAME = os.environ["CLUSTER_NAME"]

ALARM_EMOJIS = {
    "ALARM": ":rotating_light:",
    "OK": ":white_check_mark:",
    "INSUFFICIENT_DATA": ":warning:",
}


def handler(event, context):
    for record in event["Records"]:
        message = json.loads(record["Sns"]["Message"])

        alarm_name = message.get("AlarmName", "Unknown")
        state = message.get("NewStateValue", "UNKNOWN")
        reason = message.get("NewStateReason", "")
        emoji = ALARM_EMOJIS.get(state, ":question:")

        color = "#cc0000" if state == "ALARM" else "#36a64f"

        slack_payload = {
            "channel": "#cloud-costs",
            "username": f"EKS Cost Monitor ({CLUSTER_NAME})",
            "icon_emoji": ":chart_with_downwards_trend:",
            "attachments": [
                {
                    "color": color,
                    "blocks": [
                        {
                            "type": "header",
                            "text": {
                                "type": "plain_text",
                                "text": f"{emoji} {alarm_name}",
                            },
                        },
                        {
                            "type": "section",
                            "fields": [
                                {
                                    "type": "mrkdwn",
                                    "text": f"*State:*\n{state}",
                                },
                                {
                                    "type": "mrkdwn",
                                    "text": f"*Cluster:*\n{CLUSTER_NAME}",
                                },
                            ],
                        },
                        {
                            "type": "section",
                            "text": {
                                "type": "mrkdwn",
                                "text": f"*Reason:*\n{reason}",
                            },
                        },
                    ],
                }
            ],
        }

        req = request.Request(
            SLACK_WEBHOOK_URL,
            data=json.dumps(slack_payload).encode("utf-8"),
            headers={"Content-Type": "application/json"},
            method="POST",
        )

        try:
            request.urlopen(req)
        except error.HTTPError as e:
            print(f"Slack webhook failed: {e.code} {e.read().decode()}")
            raise

    return {"statusCode": 200}
Enter fullscreen mode Exit fullscreen mode

Use the module:

# environments/production/cost-monitoring.tf
module "eks_cost_alarms" {
  source = "../../modules/eks-cost-alarms"

  cluster_name      = "prod-eks-cluster"
  slack_webhook_url = var.slack_webhook_url

  efficiency_threshold = 60
  spot_ratio_threshold = 50

  namespace_budgets = {
    "payments"      = 1600
    "platform"      = 1400
    "data-pipeline" = 900
    "ml-inference"  = 700
    "frontend"      = 400
  }
}
Enter fullscreen mode Exit fullscreen mode

Layer 4: Automated Weekly Cost Reports

This is the piece that changed team behavior the most. Every Monday at 9 AM, a CronJob posts a cost report to #cloud-costs in Slack. It shows each team's spend, week-over-week change with trend arrows, and calls out the top 3 wasters and top 3 improvers by name.

This is the expanded version of the cost_report.py script from post #7, now with week-over-week comparison and Slack formatting:

#!/usr/bin/env python3
"""
Weekly EKS cost report — posts to Slack every Monday.
Pulls data from Kubecost API, compares week-over-week, and highlights
top wasters and improvers.
"""
import json
import os
from datetime import datetime
from urllib import request, error, parse


KUBECOST_ENDPOINT = os.environ.get(
    "KUBECOST_ENDPOINT", "http://kubecost-cost-analyzer.monitoring:9090"
)
SLACK_WEBHOOK_URL = os.environ["SLACK_WEBHOOK_URL"]
CLUSTER_NAME = os.environ.get("CLUSTER_NAME", "prod-eks-cluster")

EXCLUDED_NAMESPACES = {"kube-system", "monitoring", "istio-system", "kube-node-lease"}


def fetch_kubecost_data(window):
    """Fetch cost allocation data from Kubecost API."""
    url = (
        f"{KUBECOST_ENDPOINT}/model/allocation"
        f"?window={window}&aggregate=label:team&accumulate=true"
    )
    req = request.Request(url)
    with request.urlopen(req, timeout=30) as resp:
        return json.loads(resp.read().decode())


def get_team_costs(window):
    """Extract per-team costs from Kubecost response."""
    data = fetch_kubecost_data(window)
    teams = {}
    for entry in data.get("data", []):
        for team_name, metrics in entry.items():
            if team_name in ("__idle__", "__unallocated__"):
                continue
            teams[team_name] = {
                "total": round(metrics.get("totalCost", 0), 2),
                "cpu_cost": round(metrics.get("cpuCost", 0), 2),
                "ram_cost": round(metrics.get("ramCost", 0), 2),
                "cpu_eff": round(metrics.get("cpuEfficiency", 0) * 100, 1),
                "ram_eff": round(metrics.get("ramEfficiency", 0) * 100, 1),
            }
    return teams


def build_report():
    """Build the weekly cost comparison report."""
    current = get_team_costs("7d")
    previous = get_team_costs("14d,7d")  # Previous 7-day window

    rows = []
    for team in sorted(current.keys()):
        curr_cost = current[team]["total"]
        prev_cost = previous.get(team, {}).get("total", curr_cost)
        cpu_eff = current[team]["cpu_eff"]
        ram_eff = current[team]["ram_eff"]

        if prev_cost > 0:
            change_pct = ((curr_cost - prev_cost) / prev_cost) * 100
        else:
            change_pct = 0

        if change_pct > 5:
            trend = ":arrow_up_small: :red_circle:"
        elif change_pct < -5:
            trend = ":arrow_down_small: :large_green_circle:"
        else:
            trend = ":left_right_arrow:"

        rows.append({
            "team": team,
            "current": curr_cost,
            "previous": prev_cost,
            "change_pct": change_pct,
            "trend": trend,
            "cpu_eff": cpu_eff,
            "ram_eff": ram_eff,
        })

    # Sort by absolute change to find biggest movers
    rows.sort(key=lambda r: r["change_pct"], reverse=True)
    wasters = [r for r in rows if r["change_pct"] > 5][:3]
    improvers = [r for r in rows if r["change_pct"] < -5][:3]

    # Sort final table by cost descending
    rows.sort(key=lambda r: r["current"], reverse=True)

    total_current = sum(r["current"] for r in rows)
    total_previous = sum(r.get("previous", 0) for r in rows)
    total_change = ((total_current - total_previous) / total_previous * 100) if total_previous else 0

    return rows, wasters, improvers, total_current, total_change


def format_slack_message(rows, wasters, improvers, total, total_change):
    """Format the report as Slack blocks."""
    today = datetime.now().strftime("%B %d, %Y")
    trend_emoji = ":large_green_circle:" if total_change <= 0 else ":red_circle:"

    # Build the cost table
    table_lines = ["```

"]
    table_lines.append(f"{'Team':<16} {'This Week':>10} {'Last Week':>10} {'Change':>8}  {'CPU%':>5} {'RAM%':>5}")
    table_lines.append("─" * 68)

    for r in rows:
        change_str = f"{r['change_pct']:+.1f}%"
        table_lines.append(
            f"{r['team']:<16} ${r['current']:>8,.2f} ${r['previous']:>8,.2f} "
            f"{change_str:>8}  {r['cpu_eff']:>4.0f}% {r['ram_eff']:>4.0f}%"
        )

    table_lines.append("─" * 68)
    table_lines.append(
        f"{'TOTAL':<16} ${total:>8,.2f}                {total_change:+.1f}%"
    )
    table_lines.append("

```")

    blocks = [
        {
            "type": "header",
            "text": {
                "type": "plain_text",
                "text": f":bar_chart: Weekly EKS Cost Report — {today}",
            },
        },
        {
            "type": "section",
            "text": {
                "type": "mrkdwn",
                "text": (
                    f"*Cluster:* `{CLUSTER_NAME}` | "
                    f"*Weekly Spend:* ${total:,.2f} ({total_change:+.1f}% WoW) {trend_emoji}"
                ),
            },
        },
        {
            "type": "section",
            "text": {
                "type": "mrkdwn",
                "text": "\n".join(table_lines),
            },
        },
    ]

    # Top wasters callout
    if wasters:
        waster_lines = [f"• *{w['team']}*: +{w['change_pct']:.1f}% (${w['current']:,.2f})" for w in wasters]
        blocks.append({
            "type": "section",
            "text": {
                "type": "mrkdwn",
                "text": ":rotating_light: *Top Wasters (cost increase)*\n" + "\n".join(waster_lines),
            },
        })

    # Top improvers callout
    if improvers:
        improver_lines = [f"• *{i['team']}*: {i['change_pct']:.1f}% (${i['current']:,.2f})" for i in improvers]
        blocks.append({
            "type": "section",
            "text": {
                "type": "mrkdwn",
                "text": ":trophy: *Top Improvers (cost reduction)*\n" + "\n".join(improver_lines),
            },
        })

    blocks.append({
        "type": "context",
        "elements": [
            {
                "type": "mrkdwn",
                "text": (
                    f"Data source: Kubecost | Cluster: {CLUSTER_NAME} | "
                    "<https://grafana.internal/d/eks-cost-overview|View Dashboard>"
                ),
            }
        ],
    })

    return blocks


def post_to_slack(blocks):
    """Send the formatted report to Slack."""
    payload = {
        "channel": "#cloud-costs",
        "username": f"EKS Cost Report ({CLUSTER_NAME})",
        "icon_emoji": ":bar_chart:",
        "blocks": blocks,
    }

    req = request.Request(
        SLACK_WEBHOOK_URL,
        data=json.dumps(payload).encode("utf-8"),
        headers={"Content-Type": "application/json"},
        method="POST",
    )

    try:
        request.urlopen(req)
        print("Report posted to #cloud-costs")
    except error.HTTPError as e:
        print(f"Slack post failed: {e.code} {e.read().decode()}")
        raise


def main():
    rows, wasters, improvers, total, total_change = build_report()
    blocks = format_slack_message(rows, wasters, improvers, total, total_change)
    post_to_slack(blocks)


if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

Deploy it as a Kubernetes CronJob:

# weekly-cost-report-cronjob.yaml
apiVersion: batch/v1
kind: CronJob
metadata:
  name: weekly-cost-report
  namespace: monitoring
spec:
  schedule: "0 9 * * 1"  # Every Monday at 9 AM UTC
  concurrencyPolicy: Forbid
  successfulJobsHistoryLimit: 4
  failedJobsHistoryLimit: 2
  jobTemplate:
    spec:
      backoffLimit: 2
      activeDeadlineSeconds: 300
      template:
        metadata:
          labels:
            app: weekly-cost-report
            team: platform
            cost-center: engineering
        spec:
          restartPolicy: OnFailure
          serviceAccountName: cost-reporter
          containers:
            - name: report
              image: python:3.12-slim
              command: ["python3", "/scripts/weekly_cost_report.py"]
              env:
                - name: KUBECOST_ENDPOINT
                  value: "http://kubecost-cost-analyzer.monitoring:9090"
                - name: CLUSTER_NAME
                  value: "prod-eks-cluster"
                - name: SLACK_WEBHOOK_URL
                  valueFrom:
                    secretKeyRef:
                      name: slack-webhook
                      key: url
              volumeMounts:
                - name: scripts
                  mountPath: /scripts
              resources:
                requests:
                  cpu: 50m
                  memory: 64Mi
                limits:
                  cpu: 200m
                  memory: 128Mi
          volumes:
            - name: scripts
              configMap:
                name: weekly-cost-report-script
---
apiVersion: v1
kind: ConfigMap
metadata:
  name: weekly-cost-report-script
  namespace: monitoring
data:
  weekly_cost_report.py: |
    # (paste the Python script above)
---
apiVersion: v1
kind: Secret
metadata:
  name: slack-webhook
  namespace: monitoring
type: Opaque
stringData:
  url: "https://hooks.slack.com/services/YOUR/WEBHOOK/URL"
Enter fullscreen mode Exit fullscreen mode

Here's what the Slack report looks like every Monday:

┌──────────────────────────────────────────────────────────────────┐
│  📊 Weekly EKS Cost Report — September 30, 2026                 │
│                                                                  │
│  Cluster: prod-eks-cluster                                       │
│  Weekly Spend: $878.50 (-3.2% WoW) 🟢                           │
│                                                                  │
│  Team             This Week  Last Week   Change  CPU%  RAM%      │
│  ──────────────────────────────────────────────────────────────  │
│  payments           $298.40    $387.20   -22.9%   64%   61%      │
│  platform           $274.80    $268.10    +2.5%   73%   70%      │
│  data-pipeline      $148.20    $152.40    -2.8%   67%   62%      │
│  ml-inference       $102.30     $96.80    +5.7%   52%   48%      │
│  frontend            $54.80     $53.10    +3.2%   76%   74%      │
│  ──────────────────────────────────────────────────────────────  │
│  TOTAL              $878.50    $907.60    -3.2%                   │
│                                                                  │
│  🏆 Top Improvers                                                │
│  • payments: -22.9% ($298.40) — right-sized after last report    │
│                                                                  │
│  🚨 Top Wasters                                                  │
│  • ml-inference: +5.7% ($102.30) — new model deployment          │
│                                                                  │
│  Data source: Kubecost | View Dashboard                          │
└──────────────────────────────────────────────────────────────────┘
Enter fullscreen mode Exit fullscreen mode

The Behavioral Shift

The technical stack was the easy part. The hard part was getting 40 engineers to actually care about resource costs. Here's what moved the needle:

Making It Personal

Before the monitoring stack, cost was an ops problem. Platform team owned the bill, nobody else thought about it. The weekly report changed that by putting a number next to every team name. When the payments team saw they were spending $487/week at 20% efficiency — in front of every other team — they right-sized within 48 hours without anyone asking them to.

The Cost Champion Program

We assigned one "cost champion" per team — a senior engineer who got view access to the Grafana dashboards and was CC'd on budget alarms. Their job wasn't enforcement. It was awareness. They'd mention cost in sprint planning: "hey, our spend jumped 15% this week, probably that new worker deployment — can we review the resource requests?"

The champions met monthly for 30 minutes to share what worked. The data-pipeline team's champion shared a trick for right-sizing batch jobs that the ML team adopted and saved $40/week. This kind of cross-pollination never happened when costs were invisible.

Monthly Cost Review

We added a 15-minute cost review to the existing monthly architecture review. One slide: the efficiency trends dashboard screenshotted from Grafana. No finger-pointing — just "here's where we are, here's the trend." When the trend line was going up, teams self-corrected before the next meeting. Nobody wanted to be the team that broke the streak.

What Didn't Work

  • Automated enforcement (rejecting deployments above a cost threshold): Engineers gamed it by splitting workloads across namespaces. We scrapped this after two weeks.
  • Daily reports: Too frequent. People stopped reading them by day 3. Weekly was the sweet spot — frequent enough to catch problems, infrequent enough to feel meaningful.
  • Shame-based messaging: The first version of the report highlighted "worst teams." We changed it to "top improvers" within a week. Positive reinforcement drove 3x more action than negative.

Results and ROI

Six months after deploying the monitoring stack, here's where we landed:

┌──────────────────────────────────────────────────────────────────┐
│                    6-MONTH RESULTS                                │
├──────────────────────────────┬──────────┬──────────┬─────────────┤
│ Metric                       │ Month 0  │ Month 6  │ Change      │
├──────────────────────────────┼──────────┼──────────┼─────────────┤
│ Monthly EKS spend            │ $6,200   │ $3,800   │ -39%        │
│ Waste (requested - used)     │ 52%      │ 14%      │ -73%        │
│ Teams self-correcting        │ 0 / 5    │ 5 / 5    │ 100%        │
│ Avg time to detect waste     │ ~90 days │ < 24 hrs │ -99%        │
│ VPA recommendations adopted  │ 12%      │ 78%      │ +550%       │
│ Spot instance ratio          │ 42%      │ 68%      │ +62%        │
│ Budget alarm triggers        │ —        │ 7        │ All resolved│
│ Production incidents         │ 0        │ 0        │ No change   │
└──────────────────────────────┴──────────┴──────────┴─────────────┘
Enter fullscreen mode Exit fullscreen mode

Savings Breakdown

How the $2,400/month savings broke down:

  Team self-corrections (post-report)      $1,100/mo  (46%)
  ├── Payments team right-sizing            $480
  ├── ML team batch job optimization        $340
  └── Data pipeline schedule tuning         $280

  Alarm-triggered fixes                      $720/mo  (30%)
  ├── Spot ratio recovery (2 incidents)     $380
  ├── Namespace budget catches              $220
  └── Efficiency drops caught early         $120

  Proactive optimization (dashboards)        $580/mo  (24%)
  ├── Identified zombie deployments         $310
  └── Right-sized based on trend data       $270

  ─────────────────────────────────────────────────────
  Total monthly savings                    $2,400/mo
  Monitoring stack cost                       $52/mo
  Net savings                              $2,348/mo
  ROI                                          46:1
Enter fullscreen mode Exit fullscreen mode

Monitoring Stack Cost Breakdown

  Kubecost (open-source)                     $0/mo
  Prometheus (additional storage for metrics) $18/mo  (50GB retention)
  Grafana (t3.small instance)                $15/mo
  CloudWatch alarms (3 alarms)                $3/mo
  Lambda invocations (~100/mo)                $0/mo   (free tier)
  CronJob resources (50m CPU, 64Mi)          $16/mo
  ─────────────────────────────────────────────────
  Total                                      $52/mo
Enter fullscreen mode Exit fullscreen mode

Combined Program Savings

Across both optimization posts (post #7 and this one):

  Original EKS spend (pre-optimization)   $14,200/mo
  After post #7 optimizations              $4,180/mo   (-71%)
  Cost creep after 3 months                $6,200/mo   (waste returning)
  After monitoring stack (stabilized)      $3,800/mo   (-73% from original)
  ─────────────────────────────────────────────────────────────
  Total annual savings                     $124,800/yr
  Total investment (tooling + labor)        $21,100
  Payback period                            ~2 months
  First-year ROI                           492%
Enter fullscreen mode Exit fullscreen mode

The monitoring stack didn't just recover the savings from post #7 — it actually pushed costs below our original optimization target because teams found efficiencies we'd missed.

Action Items Checklist

Week 1: Foundation

  • [ ] Deploy Kubecost with Athena/CUR integration
  • [ ] Verify per-namespace cost data is accurate (compare against AWS bill)
  • [ ] Enforce cost allocation labels via OPA Gatekeeper (see post #7)
  • [ ] Test Kubecost API endpoints manually

Week 2: Visibility

  • [ ] Deploy Prometheus recording rules for cost metrics
  • [ ] Import Grafana dashboards (cluster overview, namespace detail, efficiency trends)
  • [ ] Verify dashboards populate correctly with live data
  • [ ] Share dashboard URLs with engineering leads

Week 3: Alerts

  • [ ] Deploy Terraform alarm module (efficiency, budget, spot ratio)
  • [ ] Deploy SNS → Lambda → Slack integration
  • [ ] Test each alarm by temporarily lowering thresholds
  • [ ] Set per-namespace budgets based on current baseline + 10% buffer

Week 4: Reports + Culture

  • [ ] Deploy weekly cost report CronJob
  • [ ] Verify Slack report posts correctly on Monday
  • [ ] Assign cost champions per team (one senior engineer each)
  • [ ] Schedule first monthly cost review (15 min in existing architecture review)

Ongoing

  • [ ] Review and adjust namespace budgets quarterly
  • [ ] Update Kubecost custom pricing when AWS rates change
  • [ ] Review alarm thresholds monthly (tighten as teams improve)
  • [ ] Archive weekly reports for trend analysis

Conclusion

We spent $52/month on monitoring tools and got $2,400/month in savings — not from automation, but from making costs visible to the people creating them. The payments team didn't need a policy to right-size their pods. They needed a Slack message showing they were spending 2.5x more than every other team at 20% efficiency.

The four layers work together: Kubecost for accurate cost data, Prometheus and Grafana for trend visibility, CloudWatch alarms for automated detection, and weekly Slack reports for accountability. Skip any layer and the system doesn't drive behavior change.

If you've already optimized your EKS cluster (right-sizing, Karpenter, spot instances), the monitoring stack is what prevents regression. Without it, you'll be re-optimizing the same cluster every quarter. With it, teams self-correct — and the savings compound.

This is Part 9 of the AWS Cost Optimization Series. Part 7 covers the six optimization strategies this monitoring stack protects. Part 8 covers Lambda vs. Fargate cost analysis.

Top comments (0)