Grafana
Grafana is an open-source platform for visualizing and exploring observability data. It doesn't store most data itself; instead, it connects to data sources — Prometheus, Loki, Elasticsearch, PostgreSQL, CloudWatch, Datadog, and dozens more — and turns queries into dashboards, alerts, and ad hoc investigations.
Grafana Labs also builds the storage backends that complete the LGTM stack: Loki for logs, Grafana for visualization, Tempo for traces, and Mimir for metrics. Together with OpenTelemetry, they form a popular open-source alternative to commercial platforms like Datadog. Grafana Cloud offers the same stack as a managed service.
TL;DR
- Grafana visualizes and alerts on data it queries from many sources; it's not a database.
- Dashboards are made of panels; each panel runs queries in the data source's language (PromQL, LogQL, SQL).
- Variables make one dashboard work for every service, environment, or region.
- Explore is for ad hoc investigation; correlations link metrics to logs and traces.
- Grafana Alerting evaluates queries and routes notifications through contact points and policies.
- Manage dashboards as code (provisioning, Terraform, Grafonnet) so they're versioned and reviewable.
Quick Example
A panel query showing request rate per service from Prometheus, driven by dashboard variables:
$env and $service are dashboard variables populated from label values:
Provisioning a Prometheus data source from a file so every environment is configured identically:
$__rate_interval automatically picks a safe window based on the scrape interval and the dashboard's time range.
Core Concepts
Data Sources
Mixing sources on one dashboard is a core strength: a panel of order counts from PostgreSQL can sit next to API latency from Prometheus.
Dashboards, Panels, and Visualizations
Panels include time series, stat, gauge, bar chart, table, heatmap (ideal for latency histograms), logs, traces, and node graphs. Transformations join, filter, and reshape query results inside Grafana. Rows group panels and can repeat per variable value.
Variables and Templating
Variables (query, custom, interval, data source) turn a dashboard into a template. One "Service Overview" dashboard with a $service variable replaces fifty copies that drift apart.
Explore and Correlations
Explore is a query-first view for investigating without building a dashboard. With exemplars, trace-to-logs, and logs-to-traces configured, you can click from a latency spike to a sample trace, then to that request's logs — the heart of observability debugging.
Alerting
Grafana Alerting evaluates alert rules (queries plus thresholds or conditions) on a schedule. Contact points define where notifications go (Slack, PagerDuty, email, webhooks); notification policies route alerts by labels; silences and mute timings suppress noise during maintenance. It can also manage Prometheus-compatible rules in Mimir or Loki.
The LGTM Stack
Grafana Alloy (the successor to Grafana Agent) collects and forwards telemetry using OpenTelemetry-compatible pipelines.
Best Practices
Design Dashboards Top-Down
Start with user-facing health — request rate, error rate, latency (RED) or your SLOs — at the top. Put drill-down detail (per-instance, per-endpoint, resource saturation) further down or on linked dashboards.
One Dashboard Per Question
A dashboard should answer a specific question: "Is checkout healthy?" or "Why is this node saturated?" Walls of fifty unrelated panels are ignored during incidents.
Use Variables Instead of Copies
Template by environment, cluster, and service. Fewer dashboards means less drift and easier maintenance.
Manage Dashboards as Code
Store dashboard JSON in Git and deploy through provisioning, the Grafana Terraform provider, or Grafonnet/Grafana Foundation SDK. Review changes like any other code.
Link Signals Together
Configure exemplars, derived fields, and trace-to-logs so on-call engineers can move from symptom to cause in a few clicks.
Annotate Deploys and Incidents
Deployment annotations on time-series panels make "what changed?" obvious.
Common Mistakes
Averages Instead of Percentiles
Average latency hides the slow tail users feel. Plot p50, p95, and p99, or a heatmap of the histogram. See Tail Latency.
Alerting on Every Panel
Dashboards are for looking; alerts are for waking people up. Alert only on actionable, user-impacting conditions.
Hand-Edited Production Dashboards
Click-edited dashboards drift, break silently, and disappear when someone deletes them. Provision from code.
Expensive Queries on Shared Dashboards
A panel that scans 30 days of high-cardinality data on every refresh can overload the backend. Use recording rules, reasonable default time ranges, and longer refresh intervals.
Inconsistent Units
Unlabeled axes and mixed milliseconds/seconds cause misreads during incidents. Set units on every panel.
FAQ
What is Grafana used for?
Grafana visualizes metrics, logs, traces, and business data from many sources in dashboards, supports ad hoc exploration, and sends alerts. It's a central UI for monitoring and observability.
What's the difference between Grafana and Prometheus?
Prometheus collects and stores metrics and evaluates alert rules. Grafana queries Prometheus (and other sources) to display dashboards and can also run its own alerting. They're complementary.
Is Grafana free?
Grafana OSS is free and open source under the AGPLv3 license. Grafana Enterprise and Grafana Cloud add features, support, and managed hosting; Grafana Cloud has a free tier.
What is the LGTM stack?
Loki (logs), Grafana (visualization), Tempo (traces), and Mimir (metrics) — Grafana Labs' open-source observability stack, commonly fed by OpenTelemetry.
Should I use Grafana alerting or Prometheus Alertmanager?
Either works. Alertmanager is the native choice in Prometheus-centric setups; Grafana Alerting is convenient when alerts span multiple data sources or you want one UI for rules and notifications.
Related Topics
- Prometheus — The most common Grafana data source
- Log Aggregation — Centralizing logs, including with Loki
- Distributed Tracing — Traces viewed through Tempo
- Alerting & On-Call — Designing alerts worth waking up for
- Datadog — A commercial all-in-one alternative