Datadog
Datadog is a SaaS observability and security platform that brings infrastructure metrics, application performance monitoring (APM), distributed tracing, logs, real user monitoring, synthetic tests, profiling, and security signals into one product. Instead of assembling Prometheus, Loki, Tempo, and Grafana yourself, you install the Datadog Agent, enable integrations, and get correlated data across all signals in one UI.
The trade-off is cost and lock-in. Datadog prices by host, ingested and indexed log volume, custom metrics, and many other dimensions, and bills can grow faster than infrastructure. Teams that get the most from Datadog pair its convenience with deliberate tagging, sampling, and retention decisions.
TL;DR
- Install the Datadog Agent on hosts or as a DaemonSet; turn on integrations for your stack.
- Use unified service tagging (
env,service,version) everywhere so metrics, traces, and logs line up. - APM auto-instruments common frameworks and shows service maps, latency breakdowns, and traces.
- Logs: ingest broadly, index selectively, and use archives and rehydration for the rest.
- Monitors alert on metrics, logs, traces, and SLO burn rate; route them with tags.
- Watch custom metric cardinality and log indexing — they drive most surprise bills.
Quick Example
Deploy the agent on Kubernetes with the Helm chart, enabling APM, logs, and OpenTelemetry ingestion:
Tag the workload so every signal carries the same identity:
A metric monitor on error rate, defined as code with Terraform:
Core Concepts
The Agent and Integrations
The Agent runs on each host (or node) and collects system metrics, container and Kubernetes metadata, logs, and traces, then forwards them to Datadog. Hundreds of integrations — PostgreSQL, Redis, NGINX, Kafka, AWS, Azure, GCP — add prebuilt metrics and dashboards. Cloud integrations pull metrics directly from provider APIs without an agent.
Product Areas
Unified Service Tagging
Datadog correlates signals through tags. The reserved tags env, service, and version link a latency spike on a dashboard to the traces, logs, and deployment that caused it. Consistent tagging is the single most important Datadog configuration decision.
Metrics and Custom Metrics
Integration metrics are included with host pricing; custom metrics (your own, including those from Prometheus and OpenTelemetry) are billed per unique combination of metric name and tag values. A metric tagged with customer_id can quietly become hundreds of thousands of billable series. Metrics without Limits lets you choose which tags remain queryable.
Logs: Ingest vs Index
Datadog separates ingestion (receiving and processing) from indexing (making logs searchable for a retention period). Pipelines parse and enrich logs; exclusion filters decide what gets indexed; everything can be archived to your own object storage and rehydrated later for investigations. Logs to metrics turns high-volume logs into cheap metrics.
APM and Tracing
Datadog tracing libraries auto-instrument popular frameworks, databases, and HTTP clients, propagate trace context, and connect traces with logs via injected trace IDs. You can also send OpenTelemetry data through the Agent or directly via OTLP. Ingestion controls and retention filters decide which traces are kept.
Monitors and SLOs
Monitors evaluate metric, log, APM, anomaly, forecast, and composite conditions and notify via Slack, PagerDuty, or webhooks. SLOs track reliability targets and support burn-rate alerts. See SLOs and Alerting & On-Call.
Best Practices
Standardize Tags Before Rolling Out
Define required tags (env, service, version, team) and enforce them in deployment templates. Retrofitting tags across hundreds of services is painful.
Manage Everything as Code
Keep monitors, dashboards, SLOs, and log pipelines in Terraform or the Datadog API so they're reviewed and reproducible.
Index Logs Deliberately
Index errors and business-critical events; exclude or sample noisy debug and access logs; archive everything to cheap storage for compliance.
Control Custom Metric Cardinality
Avoid unbounded tags such as user IDs, request IDs, or full URLs. Use Metrics without Limits to drop tag combinations nobody queries.
Instrument With OpenTelemetry Where Portability Matters
Using OpenTelemetry SDKs keeps your instrumentation vendor-neutral, making a future migration or a hybrid setup much easier.
Review Usage Monthly
Use the Usage and Cost pages to spot growth by product and team, and set usage attribution tags so costs can be charged back.
Common Mistakes
Indexing Every Log
Indexing all logs by default is the fastest route to a large bill. Separate ingestion from indexing.
High-Cardinality Custom Metrics
Tagging metrics with user or order IDs turns a single metric into a huge number of billable series.
Alerting on Everything
Hundreds of noisy monitors train people to ignore pages. Alert on symptoms and SLO burn, with runbooks attached.
Inconsistent Service Names
checkout, checkout-api, and Checkout across metrics, traces, and logs break correlation.
Agents Without Resource Limits
Unbounded log collection on busy nodes can consume significant CPU and memory. Set requests and limits for the Agent.
Comparison
FAQ
What is Datadog used for?
Monitoring infrastructure, tracing application requests, managing logs, tracking user experience, and alerting on problems — all in one SaaS platform with shared tags and dashboards.
Why is Datadog so expensive?
Pricing combines many dimensions: hosts, APM hosts, indexed logs, custom metrics, and more. Costs grow with data volume and cardinality, so sampling, selective indexing, and tag discipline matter.
Does Datadog support OpenTelemetry?
Yes. The Agent can receive OTLP data, and Datadog accepts OpenTelemetry metrics, traces, and logs, so you can instrument with vendor-neutral SDKs.
Datadog or Prometheus and Grafana?
Datadog offers faster setup and a unified experience with less operational work. Prometheus and Grafana are open source and cheaper at scale but require running and integrating components yourself.
Related Topics
- Monitoring — The overall practice
- OpenTelemetry — Vendor-neutral instrumentation
- Distributed Tracing — What APM traces show
- Log Aggregation — Centralizing logs and controlling cost
- Prometheus — The open-source metrics alternative