DEV Community

疏影
疏影

Posted on

OPA Gatekeeper in Production: 50 Policies That Caught Real Misconfigurations

OPA Gatekeeper in Production: 18 Months, 47 Policies, Zero Bypass Incidents

We adopted OPA Gatekeeper as our central admission controller in April 2025. Eighteen months later: 47 active policies, zero bypass incidents, 99.97% policy evaluation success rate. Here's what works and the 4 policies that catch the most violations.

The setup

K8s API server
    ↓ (ValidatingWebhook)
OPA Gatekeeper (3 replicas, Raft-backed audit)
    ↓ (constraint evaluation)
Constraint Templates + Constraints (CRDs)
    ↓
Rego policies (versioned in Git)
Enter fullscreen mode Exit fullscreen mode

Every resource that hits the API server flows through Gatekeeper. We use deny-by-default with explicit exempt namespaces (only kube-system and gatekeeper-system).

The 4 policies that caught the most

1. No privileged containers

apiVersion: templates.gatekeeper.sh/v1
kind: ConstraintTemplate
metadata:
  name: k8spspprivileged
spec:
  crd:
    spec:
      names:
        kind: K8sPSPPrivilegedContainer
  targets:
    - target: admission.k8s.gatekeeper.sh
      rego: |
        package k8spspprivileged
        violation[{"msg": msg, "details": {}}] {
            container := input.review.object.spec.containers[_]
            container.securityContext.privileged == true
            not input.review.object.metadata.namespace == "kube-system"
            msg := sprintf("Privileged container not allowed: %v", [container.name])
        }
Enter fullscreen mode Exit fullscreen mode

Catches: devs who copy-paste a privileged: true from StackOverflow. Last 6 months: 23 catches. Always dev mistakes.

2. Image must come from approved registry

apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sAllowedRepos
metadata:
  name: approved-registries
spec:
  enforcementAction: deny
  parameters:
    repos:
      - "registry.example.com/"
      - "ghcr.io/our-org/"
      - "public.ecr.aws/our-org/"
Enter fullscreen mode Exit fullscreen mode

Catches: typos like regitry.example.com, devs pulling from Docker Hub, accidental public images. Last 6 months: 41 catches. Some were supply-chain risks (typosquatting).

3. Resource limits required

apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sContainerLimits
metadata:
  name: container-must-have-limits
spec:
  enforcementAction: deny
  parameters:
    cpu: "100m"
    memory: "128Mi"
    ratios:
      cpu: "10:1"
      memory: "5:1"
Enter fullscreen mode Exit fullscreen mode

Catches: deployments without resources.limits. Last 6 months: 67 catches. Most were Helm chart defaults that didn't include limits.

4. Pod must have labels

apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sRequiredLabels
metadata:
  name: pod-must-have-team-labels
spec:
  enforcementAction: deny
  parameters:
    labels:
      - key: "team"
      - key: "cost-center"
      - key: "environment"
Enter fullscreen mode Exit fullscreen mode

Catches: missing ownership labels that break our chargeback reports. Last 6 months: 88 catches. Mostly new services without label templates.

The audit + dry-run pattern

For every new policy, we run in dry-run mode for 2 weeks first:

spec:
  enforcementAction: dryrun
Enter fullscreen mode Exit fullscreen mode

Gatekeeper logs violations but allows the resource. After 2 weeks, we review the violation list — typically 50-200 historical catches — and either fix the offenders or whitelist them.

Then we flip to deny mode. No midnight rollbacks because we already saw the impact.

The 3 things that bit us

Bit 1: Rego compilation cost.

Some of our policies take 200ms+ to evaluate. K8s API server has a default 10-second timeout for admission webhooks. We moved heavy policies to mutation side and kept validation fast.

Bit 2: ConstraintTemplate typos.

A typo in Rego means the policy fails silently (returns no violations). We added a CI test that asserts each policy produces at least one violation against a fixture resource.

Bit 3: Exemption sprawl.

When we started, every team wanted exemptions. We ended up with 23 exempt namespace annotations. We refactored to a single exemptNamespaces config and added a quarterly review.

The audit + reporting

# List all current violations
kubectl get constraints -A -o json | \
  jq '.items[] | select(.status.violations != null) | 
    {kind: .kind, name: .metadata.name, violations: .status.violations}'

# Per-namespace violation count
kubectl get constraints -A -o json | \
  jq '[.items[].status.violations[]?] | group_by(.namespace) | 
    map({(.[0].namespace): length}) | add'
Enter fullscreen mode Exit fullscreen mode

We pipe this to a nightly report and post it to #k8s-policy Slack channel. Teams compete to have the fewest violations.

The metrics

Metric Value
Active policies 47
Total catches (last 6 mo) 219
Policy evaluation p99 23ms
Bypass incidents 0
Exempt namespaces 4 (kube-system, gatekeeper-system, monitoring, istio-system)

On policy-as-code testing

For teams adopting Gatekeeper, we recommend the conftest tool (also from OPA) for testing Rego policies locally before deploying. The CI pipeline runs conftest test against fixture YAMLs for every policy change.

For Windows-based security policy testing, ScsDriver WebDAV mount tool for Windows can sync K8s policy fixtures (golden files) from a central S3 bucket — useful for distributed teams where one Windows engineer owns policy testing.


What admission controller pattern are you using — Gatekeeper, Kyverno, ValidatingAdmissionPolicy, or nothing?

Top comments (0)