OpenTelemetry
OpenTelemetry (OTel) is an open-source, vendor-neutral standard for generating, collecting, and exporting telemetry — traces, metrics, and logs (with profiles emerging). It's a CNCF project formed from the merger of OpenTracing and OpenCensus, and it's supported by every major observability vendor, including Datadog, Grafana, New Relic, Honeycomb, Elastic, and the big cloud providers.
The key idea: instrument once, send anywhere. Your code and libraries emit telemetry through the OpenTelemetry API; the SDK and the Collector decide where it goes. Switching or combining backends becomes configuration, not a re-instrumentation project.
TL;DR
- OTel standardizes APIs, SDKs, a wire protocol (OTLP), and semantic conventions for telemetry.
- Auto-instrumentation covers common frameworks and libraries; add manual spans for business logic.
- Context propagation (W3C
traceparent) links spans across services into one trace. - The Collector receives, processes (batching, filtering, redaction, sampling), and exports telemetry.
- Use semantic conventions for attribute names so data is consistent across services and tools.
- Control volume with head sampling in SDKs and tail sampling in the Collector.
Quick Example
Instrument a Node.js service: auto-instrumentation plus a custom span with attributes.
A Collector that receives OTLP, batches and redacts, tail-samples, and fans out to two backends:
Core Concepts
Signals
API, SDK, and Instrumentation
The API is what libraries and applications call; it's a safe no-op if no SDK is configured. The SDK implements sampling, processing, and export. Instrumentation libraries (and zero-code agents for Java, .NET, Python, Node.js) hook into frameworks, HTTP clients, and database drivers automatically.
Traces, Spans, and Context Propagation
A trace is a tree of spans, each with a name, timing, attributes, events, and status. Context propagation passes the trace ID and parent span ID between services — in HTTP via the W3C traceparent header, in messaging via message headers — so work across services joins one trace. Baggage carries additional key-value context. See Distributed Tracing.
Resources and Semantic Conventions
A resource describes what's producing telemetry (service.name, service.version, deployment.environment.name, k8s.pod.name). Semantic conventions define standard attribute names (http.request.method, http.response.status_code, db.system), making dashboards and queries portable across services and vendors.
OTLP and the Collector
OTLP is OpenTelemetry's protocol over gRPC or HTTP. The Collector is a standalone service built from receivers (OTLP, Prometheus scrape, Kafka, host metrics, filelog), processors (batch, memory limiter, attributes, filter, tail sampling, k8s attributes), and exporters (OTLP, Prometheus, vendor-specific). Deploy it as an agent (per node or sidecar) and/or a gateway (central pool).
Sampling
- Head sampling decides at the start of a trace (e.g. keep 10%), cheap but blind to outcome.
- Tail sampling decides after the trace completes in the Collector, so you can keep all errors and slow requests while sampling routine traffic — at the cost of buffering.
Best Practices
Start With Auto-Instrumentation
It gives immediate value for HTTP, database, and messaging calls. Add manual spans and attributes for business operations afterward.
Always Set service.name and Environment
Unnamed services show up as unknown_service and break service maps and dashboards.
Route Through a Collector
Applications export to a local Collector; the Collector handles retries, batching, redaction, and backend changes without redeploying apps.
Follow Semantic Conventions
Use standard attribute names and consistent units so cross-service queries work.
Redact Sensitive Data
Strip PII and secrets from attributes and logs in the Collector before data leaves your network.
Correlate Logs With Traces
Use OTel logging bridges or inject trace_id and span_id into logs so you can pivot between signals.
Common Mistakes
Broken Context Propagation
Missing propagators, custom HTTP clients, or message queues without header propagation split one request into many disconnected traces.
High-Cardinality Metric Attributes
User IDs or full URLs on metrics explode storage costs. Keep those on spans, not metrics.
Forgetting to End Spans
Spans that never end are never exported and can leak memory. Use finally blocks or helper wrappers.
Exporting Directly to a Vendor From Every Service
It works, but changing vendors or adding redaction later means touching every service.
100% Trace Sampling at Scale
Keeping every trace in high-traffic systems is expensive and rarely useful. Use tail sampling to keep what matters.
Comparison
FAQ
What is OpenTelemetry?
An open standard and set of tools for instrumenting software to produce traces, metrics, and logs, and for collecting and exporting that telemetry to any compatible backend.
Is OpenTelemetry a monitoring tool?
No. It produces and transports telemetry. You still need backends to store and visualize it, such as Grafana Tempo and Prometheus, Jaeger, or a commercial platform.
Do I need the OpenTelemetry Collector?
Not strictly — SDKs can export directly. But the Collector is recommended for production because it centralizes batching, retries, sampling, redaction, and routing.
What's the difference between head and tail sampling?
Head sampling decides whether to keep a trace when it starts. Tail sampling decides after it completes, allowing policies like "keep all errors and slow traces."
Does OpenTelemetry replace Prometheus?
Not necessarily. OTel can produce metrics and export them to Prometheus-compatible storage; the Collector can also scrape Prometheus endpoints. They work together.
Related Topics
- Distributed Tracing — What traces reveal and how to use them
- Metrics & Time Series — Designing metrics that scale
- Logging — Structured logs correlated with traces
- Prometheus — A common metrics backend
- Datadog — A commercial backend that accepts OTLP