Prometheus Exporters & Service Discovery
Prometheus uses a pull model: it periodically scrapes HTTP endpoints that expose metrics in its text format. Your own services expose metrics through client libraries (see metric types), but most infrastructure (Linux hosts, databases, load balancers, message brokers) doesn't speak Prometheus natively. Exporters fill that gap: small processes that translate a system's internal stats into Prometheus metrics.
The other half is service discovery: in dynamic environments like Kubernetes and cloud autoscaling groups, targets appear and disappear constantly. Prometheus discovers them automatically, and relabeling shapes which targets are scraped and what labels they carry.
TL;DR
- Exporters expose metrics for third-party systems:
node_exporter(hosts),blackbox_exporter(probes),postgres_exporter,redis_exporter,mysqld_exporter,kafka_exporter, and many more. - A scrape config defines targets, interval, timeout, path, and auth for a job.
- Service discovery finds targets dynamically (Kubernetes, EC2, Consul, DNS, file-based, HTTP).
relabel_configsfilter and rename targets before scraping;metric_relabel_configsfilter series after scraping.- On Kubernetes, the Prometheus Operator uses ServiceMonitor and PodMonitor resources instead of hand-written scrape configs.
- Use Pushgateway only for short-lived batch jobs, never as a general push pipeline.
Quick Example
A scrape configuration combining static, blackbox, and Kubernetes discovery:
Core Concepts
Exporters
Many modern systems expose Prometheus metrics natively (Kubernetes components, etcd, Traefik, Envoy, CockroachDB, Grafana), so no exporter is needed. Exporters run as sidecars, DaemonSets (node_exporter), or standalone services.
Scrape Configs and Jobs
A job groups targets serving the same purpose. Each scraped series gets job and instance labels automatically, plus an up series (1 if the scrape succeeded, 0 if not), the most basic health signal. Per-job options include scrape_interval, scrape_timeout, metrics_path, scheme, TLS and auth settings, honor_labels, and sample_limit to guard against cardinality explosions.
Service Discovery
Discovered targets carry __meta_* labels (pod name, namespace, instance tags) available during relabeling.
Relabeling
Relabeling is a sequence of rules applied to label sets:
relabel_configs(before scraping): keep or drop targets, set__address__,__metrics_path__, and__scheme__, and copy__meta_*values into permanent labels.metric_relabel_configs(after scraping, before storage): drop expensive metrics or labels, and rename series.
Common actions are keep, drop, replace, labelmap, labeldrop, labelkeep, and hashmod (for sharding scrapes across Prometheus instances). Labels starting with __ are dropped after relabeling.
Prometheus Operator
On Kubernetes, the Prometheus Operator (bundled in kube-prometheus-stack) manages Prometheus and Alertmanager as custom resources. Teams declare what to scrape alongside their apps:
PodMonitor, Probe (blackbox), and PrometheusRule resources cover pods, probes, and alert rules. That makes monitoring self-service and GitOps-friendly.
Pushgateway
Short-lived batch jobs may finish before Prometheus scrapes them. They can push final metrics (last success time, duration, records processed) to a Pushgateway, which Prometheus scrapes. It's not a general push mechanism: pushed metrics never expire automatically, and the gateway becomes a single point of failure. For long-running services, always expose /metrics. OpenTelemetry's OTLP push into Prometheus 3.x is a separate option for push-based pipelines.
Best Practices
Probe From the Outside Too
Internal metrics can look healthy while users can't reach you (DNS, TLS, CDN, or firewall issues). Blackbox probes against public endpoints, from more than one location, catch those failures, and TLS expiry checks prevent certificate outages.
Standardize Labels at Discovery Time
Map discovery metadata to consistent labels (namespace, pod, service, env, team) in relabeling, so every dashboard and alert can filter the same way across jobs.
Drop What You Don't Use
Exporters often expose hundreds of metrics per target. Drop unused high-cardinality metrics with metric_relabel_configs, which saves memory and storage and speeds up queries. See Prometheus long-term storage.
Secure Exporter Endpoints
Metrics can leak internal details: hostnames, versions, and query text in database exporters. Bind exporters to internal networks, use TLS and basic auth (supported by the exporter toolkit), and grant database exporters read-only monitoring roles.
Common Mistakes
Using Pushgateway for Services
Pushing metrics from long-running services through Pushgateway loses the up signal, keeps stale series forever after an instance dies, and centralizes failure. Let Prometheus scrape the service.
Scraping Through a Load Balancer
Scraping api.internal:8080/metrics behind a load balancer hits a random instance each time, producing jumbled counters. Scrape each instance directly via service discovery.
Missing Monitoring for Scrape Failures
If a target silently stops being scraped, its alerts can't fire. Alert on up == 0 for important jobs, and on scrape-duration and sample-limit breaches.
FAQ
Do I need an exporter for my application?
No, if you can instrument it: add a Prometheus or OpenTelemetry client library and expose /metrics. Exporters are for software you don't control, such as databases, operating systems, and network gear.
How do I monitor Kubernetes with Prometheus?
The standard path is the kube-prometheus-stack Helm chart: Prometheus Operator, Prometheus, Alertmanager, Grafana, node_exporter, kube-state-metrics, and default dashboards and alerts for cluster health. Add ServiceMonitors for your own services.
What's the difference between relabel_configs and metric_relabel_configs?
relabel_configs runs on targets before scraping. It decides whether and how to scrape each target, and sets target labels. metric_relabel_configs runs on each scraped series before storage, to drop or modify series and labels.
How often should Prometheus scrape?
15–60 seconds is typical, with 30s a common default. Shorter intervals give finer resolution but multiply storage and load. Keep the interval consistent within a job, and make the rate() windows in your queries at least 4× the interval.
Related Topics
- Prometheus — The monitoring system overview
- Prometheus Metric Types — Instrumenting your own services
- Prometheus Long-Term Storage — Scaling beyond one server
- Kubernetes — Service discovery for pods and services
- Monitoring — Monitoring strategy in general
- Grafana — Dashboards for exporter metrics