PromQL
PromQL (Prometheus Query Language) is how you get answers out of Prometheus: request rates, error ratios, latency percentiles, saturation, and anything else you can compute from time series. The same queries power Grafana dashboards, alerting rules, recording rules, and ad hoc debugging during incidents, and PromQL-compatible backends (Thanos, Mimir, VictoriaMetrics, cloud managed Prometheus) accept it too.
PromQL is compact and functional: you select series by metric name and labels, transform them with functions like rate(), and aggregate across labels with operators like sum by. A handful of patterns (rate of counters, ratios, histogram_quantile) cover most real-world queries, and a few gotchas (rate-then-sum, counter types, label matching) cause most wrong answers.
TL;DR
- Select series with
metric{label="value"}; matchers support=,!=,=~(regex), and!~. - An instant vector has one sample per series; a range vector (
[5m]) has a window of samples. - For counters, use
rate()(per-second average) orincrease()over a range. Never graph a raw counter. - Aggregate with
sum,avg,max,count,topk, and more, keeping labels withby (...)or dropping them withwithout (...). - Latency percentiles:
histogram_quantile(0.99, sum by (le) (rate(x_bucket[5m]))). - Rate first, then aggregate, and keep label sets aligned when dividing series.
Quick Example
The core RED queries (rate, errors, duration) for an HTTP service:
Core Concepts
Data Model
Every time series is a metric name plus a set of labels:
Each unique combination of labels is a separate series. Selectors filter them:
Instant Vectors, Range Vectors, and Scalars
- Instant vector: the latest sample of each matching series at the evaluation time. It's what you graph or alert on.
- Range vector:
metric[5m], all samples in the last 5 minutes per series. It's the input to functions likerate(), and can't be graphed directly. - Scalar: a single number, such as
0.99ortime().
offset 1h shifts the evaluation window back in time (rate(x[5m] offset 1d) for "same time yesterday").
Counters and rate()
Counters only go up (and reset to zero when a process restarts). Their raw values are meaningless on a graph; what matters is how fast they increase:
rate handles counter resets automatically. Choose a range of at least 4× the scrape interval (for example [2m] at 30s scrapes, [5m] is a common default). See Prometheus metric types.
Aggregation Operators
Others include avg, min, max, stddev, count_values, bottomk, and group.
Histograms and Percentiles
A histogram metric exposes cumulative _bucket{le="…"} counters, plus _sum and _count. To compute a percentile, take the rate of the buckets, aggregate while keeping le, then apply histogram_quantile:
Quantiles are estimated by interpolating within buckets, so accuracy depends on bucket boundaries. Native histograms (stable in Prometheus 3.x) use exponential buckets and improve accuracy without manual bucket design.
Binary Operators and Vector Matching
Arithmetic (+ - * / % ^), comparison (> < == !=), and logical (and, or, unless) operators work between vectors by matching series with identical label sets. When label sets differ, control matching:
Comparisons filter by default (x > 0.9 keeps only series above 0.9), which is how alert expressions work. Add bool to return 0/1 instead.
Useful Functions
absent() (alert when a metric disappears), predict_linear() (disk full in 4 hours?), delta()/deriv() (gauges), clamp_min/clamp_max, label_replace, *_over_time functions (avg_over_time, max_over_time, quantile_over_time), changes(), resets(), and time().
Subqueries apply range functions to expressions: max_over_time(rate(x[5m])[1h:1m]), the peak 5-minute rate within the last hour.
Query Patterns
Best Practices
Rate Before You Aggregate
Always apply rate() to individual series first, then sum. rate(sum(x)[5m]) isn't valid without a subquery, and summing counters first breaks counter-reset handling.
Use Recording Rules for Expensive Queries
Precompute heavy or frequently used expressions (per-service request rates, error ratios) as recording rules named level:metric:operations, for example service:http_requests:rate5m. Dashboards and alerts become fast and consistent.
Keep Cardinality in Mind
Queries touching millions of series are slow and memory-hungry. Filter with label matchers early, avoid regexes over high-cardinality labels, and don't aggregate by labels like user_id. See Prometheus long-term storage.
Match Range Windows to Scrape Intervals
Too short a range ([30s] with 30s scrapes) yields gaps and empty results. Use at least 4× the scrape interval, and consider $__rate_interval in Grafana.
Common Mistakes
Aggregating Away le Before histogram_quantile
Averaging Percentiles
avg(histogram_quantile(0.99, ...)) across instances isn't the p99 of the fleet. Aggregate the buckets first, then compute the quantile once.
Dividing Series With Mismatched Labels
rate(errors_total[5m]) / rate(requests_total[5m]) returns nothing if the two metrics carry different labels (for example status exists only on one side). Aggregate both sides to the same labels, or use on()/ignoring().
FAQ
What's the difference between rate and irate?
rate averages the per-second increase over the whole range, which is smooth and robust, so it's the right choice for alerts and most dashboards. irate uses only the last two data points, so it reacts instantly to changes but is noisy. Use it for zoomed-in debugging graphs of volatile counters.
Why does increase() return non-integer values?
increase is computed from rate with extrapolation to the window edges, so it can return fractional results even for integer counters. For exact counts, reduce the window's extrapolation effects or accept the approximation. It's designed for trends, not accounting.
How do I get a p95 across all instances?
Sum the histogram bucket rates across instances while keeping le, then apply histogram_quantile(0.95, …). With summaries (precomputed quantiles), you can't aggregate across instances meaningfully, which is one reason histograms are preferred.
Can PromQL join metrics like SQL?
Partially. Vector matching with on(), ignoring(), group_left, and group_right joins series by labels, commonly to attach metadata from "info" metrics. It's not a general relational join. Complex correlation is often easier in Grafana transformations or by adding labels at instrumentation time.
Related Topics
- Prometheus — The monitoring system overview
- Prometheus Metric Types — Counters, gauges, histograms, summaries
- Prometheus Alertmanager — Alerting on PromQL expressions
- Grafana — Visualizing PromQL queries
- SLOs — Burn-rate queries and error budgets
- Metrics — Metrics design in general