Most people set up fail2ban and forget about it. It's running, it's banning IPs, but you have no
idea what's actually happening unless you SSH in and check the logs manually.
This post shows how to add fail2ban metrics to your Prometheus stack and visualize them in Grafana
— so you see ban activity on the same dashboard as CPU, memory, and disk.
## What you'll see
- Active bans per jail in real time
- Total banned IPs over time
- Which jails are firing most — SSH? Nginx? Recidive?
- Ban spikes correlated with traffic and load
## Step 1 — Add the exporter
fail2ban doesn't expose Prometheus metrics natively. Add
fail2ban-exporter to your
docker-compose.yml:
yaml
fail2ban-exporter:
image: braedon/prometheus-fail2ban-exporter:latest
volumes:
- /var/run/fail2ban/fail2ban.sock:/var/run/fail2ban/fail2ban.sock:ro
ports:
- "9191:9191"
restart: unless-stopped
The exporter reads fail2ban's Unix socket and exposes /metrics on port 9191.
Step 2 — Add a scrape job to Prometheus
In prometheus.yml:
- job_name: 'fail2ban'
static_configs:
- targets: ['fail2ban-exporter:9191']
Restart Prometheus. You should now see metrics like fail2ban_banned_ips and fail2ban_enabled_jails
in the Prometheus UI.
Step 3 — Add an alert rule
The useful one — fires when more than 10 IPs get banned in 5 minutes on any jail:
- alert: Fail2BanHighBanRate
expr: increase(fail2ban_banned_ips_total[5m]) > 10
for: 0m
annotations:
summary: "High ban rate — jail {{ $labels.jail }}"
This catches coordinated brute-force attempts before they become a problem.
Step 4 — Import the Grafana dashboard
Download the fail2ban dashboard JSON from the monitoring-stack repo (grafana-dashboards/json/) and
import it in Grafana via Dashboards → Import.
The full repo also includes a one-command Ansible deployment for the entire stack — Prometheus,
Grafana, Alertmanager, Node Exporter, Blackbox Exporter — if you want to set everything up at once
rather than piece by piece.
Why bother?
The first time I saw a ban spike in Grafana — 40 IPs banned in under 2 minutes at 3am — I realized
I had zero visibility into this before. It was happening, I just couldn't see it.
Correlating ban events with CPU and network in one view makes it obvious when something is
actually an attack versus a misconfigured client.
---
Repo: github.com/Airat71/monitoring-stack

Top comments (0)