Prometheus Alerting & Alertmanager

Prometheus alerting has two parts. Alerting rules, evaluated by Prometheus, are PromQL expressions that fire when a condition holds for long enough. Alertmanager receives those firing alerts and decides what to do with them: it groups related alerts, deduplicates them across Prometheus replicas, routes them to the right team and channel, inhibits noisy downstream alerts, and supports silences during maintenance.

The technology is straightforward. The hard part is alert design: paging only on problems users feel, with enough context to act, and without waking people for noise. Good alerting follows the practices of SRE: alert on symptoms and SLO burn, not every cause.

TL;DR

Quick Example

Alerting rules (loaded by Prometheus):

Alertmanager routing:

Core Concepts

Alerting Rules

Prometheus evaluates each rule group every evaluation_interval. An alert is:

for filters out brief blips. keep_firing_for (Prometheus 2.42+) keeps an alert firing for a while after the condition clears, which prevents flapping. Labels drive routing (severity, team, service). Annotations give responders context, templated with {{ $labels.x }} and {{ $value }}.

Symptom-Based Alerting

Page on what users feel: elevated error rate, high latency, unavailability, and data freshness for pipelines. Causes (high CPU, a pod restart, disk at 70%) belong on dashboards or in low-priority tickets, unless they predict imminent user impact (disk full within hours via predict_linear).

SLO Burn-Rate Alerts

A burn rate is how fast you're consuming your error budget relative to plan. With a 99.9% SLO, the budget is 0.1% errors, and a burn rate of 14.4 exhausts a 30-day budget in about 2 days. The Google SRE multi-window, multi-burn-rate approach:

The short window makes alerts resolve quickly once the problem stops. Tools like Sloth and Pyrra generate these rules from SLO definitions.

Alertmanager Concepts

Best Practices

Make Every Page Actionable

If the responder can't do anything about it right now, it shouldn't page. Downgrade to a ticket or a dashboard, or delete it. Review pages after every on-call shift, and prune noisy alerts relentlessly. See incident management.

Include a Runbook and a Dashboard

Annotations should link to a runbook (what it means, how to diagnose, how to mitigate) and a dashboard scoped to the problem. That context turns 20 minutes of orientation into 2.

Test Alert Rules

Validate syntax with promtool check rules, and write unit tests with promtool test rules, feeding synthetic series and asserting which alerts fire when. Keep rules in version control and deploy via CI.

Monitor the Monitoring

Run an always-firing Watchdog alert routed to a dead man's switch service (such as healthchecks.io or PagerDuty's heartbeat), so you're notified if Prometheus or Alertmanager stops working. Alert on failed notifications (alertmanager_notifications_failed_total) and rule evaluation failures.

Common Mistakes

Alerting on Every Cause

Page on latency and errors instead, and investigate CPU from the dashboard when they fire.

No for Duration

Without for, a single bad scrape or momentary spike pages someone. Use for (or burn-rate windows) proportional to how quickly you need to react.

Alerts That Vanish When Data Disappears

rate(errors_total[5m]) > 1 can't fire if the service is down and not exporting metrics at all. Pair symptom alerts with up == 0 or absent() alerts for critical jobs.

FAQ

What's the difference between Prometheus alerting rules and Alertmanager?

Prometheus evaluates rules and decides whether an alert is firing. Alertmanager decides who gets notified, how, and when: grouping, routing, deduplication, inhibition, and silencing. Separating them lets multiple Prometheus servers share one notification pipeline.

How do I avoid duplicate notifications from HA Prometheus pairs?

Configure both Prometheus replicas to send alerts to the same Alertmanager cluster. Identical alerts from both are deduplicated automatically. Make sure external labels like replica are dropped (via alert_relabel_configs) so the alerts look identical.

Should I alert in Grafana or Alertmanager?

Both work. Grafana Alerting can query many data sources and has a friendly UI; Prometheus rules plus Alertmanager are file-based, version-controlled, and evaluated close to the data. Many teams keep critical SLO alerts as Prometheus rules in Git, and use Grafana alerts for cross-source or ad hoc cases.

What severity levels should we use?

Keep it simple: page (wake someone up: user impact now or imminent), ticket (needs attention within business hours), and optionally info (dashboards and chat only). More levels usually add confusion rather than clarity.

Related Topics

References