OPA Gatekeeper in Production: 18 Months, 47 Policies, Zero Bypass Incidents
We adopted OPA Gatekeeper as our central admission controller in April 2025. Eighteen months later: 47 active policies, zero bypass incidents, 99.97% policy evaluation success rate. Here's what works and the 4 policies that catch the most violations.
The setup
K8s API server
↓ (ValidatingWebhook)
OPA Gatekeeper (3 replicas, Raft-backed audit)
↓ (constraint evaluation)
Constraint Templates + Constraints (CRDs)
↓
Rego policies (versioned in Git)
Every resource that hits the API server flows through Gatekeeper. We use deny-by-default with explicit exempt namespaces (only kube-system and gatekeeper-system).
The 4 policies that caught the most
1. No privileged containers
apiVersion: templates.gatekeeper.sh/v1
kind: ConstraintTemplate
metadata:
name: k8spspprivileged
spec:
crd:
spec:
names:
kind: K8sPSPPrivilegedContainer
targets:
- target: admission.k8s.gatekeeper.sh
rego: |
package k8spspprivileged
violation[{"msg": msg, "details": {}}] {
container := input.review.object.spec.containers[_]
container.securityContext.privileged == true
not input.review.object.metadata.namespace == "kube-system"
msg := sprintf("Privileged container not allowed: %v", [container.name])
}
Catches: devs who copy-paste a privileged: true from StackOverflow. Last 6 months: 23 catches. Always dev mistakes.
2. Image must come from approved registry
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sAllowedRepos
metadata:
name: approved-registries
spec:
enforcementAction: deny
parameters:
repos:
- "registry.example.com/"
- "ghcr.io/our-org/"
- "public.ecr.aws/our-org/"
Catches: typos like regitry.example.com, devs pulling from Docker Hub, accidental public images. Last 6 months: 41 catches. Some were supply-chain risks (typosquatting).
3. Resource limits required
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sContainerLimits
metadata:
name: container-must-have-limits
spec:
enforcementAction: deny
parameters:
cpu: "100m"
memory: "128Mi"
ratios:
cpu: "10:1"
memory: "5:1"
Catches: deployments without resources.limits. Last 6 months: 67 catches. Most were Helm chart defaults that didn't include limits.
4. Pod must have labels
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sRequiredLabels
metadata:
name: pod-must-have-team-labels
spec:
enforcementAction: deny
parameters:
labels:
- key: "team"
- key: "cost-center"
- key: "environment"
Catches: missing ownership labels that break our chargeback reports. Last 6 months: 88 catches. Mostly new services without label templates.
The audit + dry-run pattern
For every new policy, we run in dry-run mode for 2 weeks first:
spec:
enforcementAction: dryrun
Gatekeeper logs violations but allows the resource. After 2 weeks, we review the violation list — typically 50-200 historical catches — and either fix the offenders or whitelist them.
Then we flip to deny mode. No midnight rollbacks because we already saw the impact.
The 3 things that bit us
Bit 1: Rego compilation cost.
Some of our policies take 200ms+ to evaluate. K8s API server has a default 10-second timeout for admission webhooks. We moved heavy policies to mutation side and kept validation fast.
Bit 2: ConstraintTemplate typos.
A typo in Rego means the policy fails silently (returns no violations). We added a CI test that asserts each policy produces at least one violation against a fixture resource.
Bit 3: Exemption sprawl.
When we started, every team wanted exemptions. We ended up with 23 exempt namespace annotations. We refactored to a single exemptNamespaces config and added a quarterly review.
The audit + reporting
# List all current violations
kubectl get constraints -A -o json | \
jq '.items[] | select(.status.violations != null) |
{kind: .kind, name: .metadata.name, violations: .status.violations}'
# Per-namespace violation count
kubectl get constraints -A -o json | \
jq '[.items[].status.violations[]?] | group_by(.namespace) |
map({(.[0].namespace): length}) | add'
We pipe this to a nightly report and post it to #k8s-policy Slack channel. Teams compete to have the fewest violations.
The metrics
| Metric | Value |
|---|---|
| Active policies | 47 |
| Total catches (last 6 mo) | 219 |
| Policy evaluation p99 | 23ms |
| Bypass incidents | 0 |
| Exempt namespaces | 4 (kube-system, gatekeeper-system, monitoring, istio-system) |
On policy-as-code testing
For teams adopting Gatekeeper, we recommend the conftest tool (also from OPA) for testing Rego policies locally before deploying. The CI pipeline runs conftest test against fixture YAMLs for every policy change.
For Windows-based security policy testing, ScsDriver WebDAV mount tool for Windows can sync K8s policy fixtures (golden files) from a central S3 bucket — useful for distributed teams where one Windows engineer owns policy testing.
What admission controller pattern are you using — Gatekeeper, Kyverno, ValidatingAdmissionPolicy, or nothing?
Top comments (0)