Grafana

Grafana is an open-source platform for visualizing and exploring observability data. It doesn't store most data itself; instead, it connects to data sources — Prometheus, Loki, Elasticsearch, PostgreSQL, CloudWatch, Datadog, and dozens more — and turns queries into dashboards, alerts, and ad hoc investigations.

Grafana Labs also builds the storage backends that complete the LGTM stack: Loki for logs, Grafana for visualization, Tempo for traces, and Mimir for metrics. Together with OpenTelemetry, they form a popular open-source alternative to commercial platforms like Datadog. Grafana Cloud offers the same stack as a managed service.

TL;DR

Quick Example

A panel query showing request rate per service from Prometheus, driven by dashboard variables:

$env and $service are dashboard variables populated from label values:

Provisioning a Prometheus data source from a file so every environment is configured identically:

$__rate_interval automatically picks a safe window based on the scrape interval and the dashboard's time range.

Core Concepts

Data Sources

Mixing sources on one dashboard is a core strength: a panel of order counts from PostgreSQL can sit next to API latency from Prometheus.

Dashboards, Panels, and Visualizations

Panels include time series, stat, gauge, bar chart, table, heatmap (ideal for latency histograms), logs, traces, and node graphs. Transformations join, filter, and reshape query results inside Grafana. Rows group panels and can repeat per variable value.

Variables and Templating

Variables (query, custom, interval, data source) turn a dashboard into a template. One "Service Overview" dashboard with a $service variable replaces fifty copies that drift apart.

Explore and Correlations

Explore is a query-first view for investigating without building a dashboard. With exemplars, trace-to-logs, and logs-to-traces configured, you can click from a latency spike to a sample trace, then to that request's logs — the heart of observability debugging.

Alerting

Grafana Alerting evaluates alert rules (queries plus thresholds or conditions) on a schedule. Contact points define where notifications go (Slack, PagerDuty, email, webhooks); notification policies route alerts by labels; silences and mute timings suppress noise during maintenance. It can also manage Prometheus-compatible rules in Mimir or Loki.

The LGTM Stack

Grafana Alloy (the successor to Grafana Agent) collects and forwards telemetry using OpenTelemetry-compatible pipelines.

Best Practices

Design Dashboards Top-Down

Start with user-facing health — request rate, error rate, latency (RED) or your SLOs — at the top. Put drill-down detail (per-instance, per-endpoint, resource saturation) further down or on linked dashboards.

One Dashboard Per Question

A dashboard should answer a specific question: "Is checkout healthy?" or "Why is this node saturated?" Walls of fifty unrelated panels are ignored during incidents.

Use Variables Instead of Copies

Template by environment, cluster, and service. Fewer dashboards means less drift and easier maintenance.

Manage Dashboards as Code

Store dashboard JSON in Git and deploy through provisioning, the Grafana Terraform provider, or Grafonnet/Grafana Foundation SDK. Review changes like any other code.

Link Signals Together

Configure exemplars, derived fields, and trace-to-logs so on-call engineers can move from symptom to cause in a few clicks.

Annotate Deploys and Incidents

Deployment annotations on time-series panels make "what changed?" obvious.

Common Mistakes

Averages Instead of Percentiles

Average latency hides the slow tail users feel. Plot p50, p95, and p99, or a heatmap of the histogram. See Tail Latency.

Alerting on Every Panel

Dashboards are for looking; alerts are for waking people up. Alert only on actionable, user-impacting conditions.

Hand-Edited Production Dashboards

Click-edited dashboards drift, break silently, and disappear when someone deletes them. Provision from code.

Expensive Queries on Shared Dashboards

A panel that scans 30 days of high-cardinality data on every refresh can overload the backend. Use recording rules, reasonable default time ranges, and longer refresh intervals.

Inconsistent Units

Unlabeled axes and mixed milliseconds/seconds cause misreads during incidents. Set units on every panel.

FAQ

What is Grafana used for?

Grafana visualizes metrics, logs, traces, and business data from many sources in dashboards, supports ad hoc exploration, and sends alerts. It's a central UI for monitoring and observability.

What's the difference between Grafana and Prometheus?

Prometheus collects and stores metrics and evaluates alert rules. Grafana queries Prometheus (and other sources) to display dashboards and can also run its own alerting. They're complementary.

Is Grafana free?

Grafana OSS is free and open source under the AGPLv3 license. Grafana Enterprise and Grafana Cloud add features, support, and managed hosting; Grafana Cloud has a free tier.

What is the LGTM stack?

Loki (logs), Grafana (visualization), Tempo (traces), and Mimir (metrics) — Grafana Labs' open-source observability stack, commonly fed by OpenTelemetry.

Should I use Grafana alerting or Prometheus Alertmanager?

Either works. Alertmanager is the native choice in Prometheus-centric setups; Grafana Alerting is convenient when alerts span multiple data sources or you want one UI for rules and notifications.

Related Topics

References