OpenTelemetry Instrumentation

Instrumentation is the code that makes an application emit telemetry. OpenTelemetry provides a vendor-neutral API and SDK in every major language, plus instrumentation libraries for popular frameworks (HTTP servers and clients, database drivers, message queues, gRPC) and zero-code agents that attach them automatically. You instrument once and send data to any backend.

The usual approach is layered. Start with auto-instrumentation to get request traces and RED metrics for free, then add manual spans and attributes around the business logic that matters: checkout steps, payment calls, cache decisions. The best instrumentation captures the context you'll need at 3 a.m. to explain why a request was slow or failed.

TL;DR

Quick Example

Python: auto-instrumentation plus custom spans and metrics.

The resulting trace shows the incoming HTTP request, place_order, the payment HTTP call, and the database queries as one tree.

Core Concepts

API vs SDK

Configure the SDK once at startup, programmatically or through OTEL_* environment variables, which are standardized across languages.

Zero-Code Instrumentation

Auto-instrumentation gives you server and client spans, DB and messaging spans, and HTTP metrics following semantic conventions.

Spans

A span records one operation: name, start and end time, kind (server, client, internal, producer, consumer), attributes, events (timestamped annotations, including exceptions), status (unset, ok, error), and links (to related traces, for example in batch processing). Spans nest via context: the active span becomes the parent of spans created inside it, and context crosses service boundaries via propagation.

Good span names are low-cardinality operation names (GET /orders/{id}, place_order), never including IDs.

Metrics Instruments

Metric attributes follow the same cardinality rules as Prometheus labels. Aggregation temporality (cumulative vs delta) depends on the backend: Prometheus expects cumulative.

Logs and Correlation

OpenTelemetry doesn't replace your logging library. Log bridges or appenders (for Log4j, Logback, Python logging, .NET ILogger, Winston or Pino) capture existing logs, attach the current trace and span IDs, and export them via OTLP. Correlated logs let you jump from a slow span to exactly the logs it produced. See logging.

Resources

A Resource describes the entity producing telemetry: service.name (required in practice), service.version, service.namespace, deployment.environment.name, plus host, container, cloud, and Kubernetes attributes added by resource detectors or the Collector's k8sattributes processor. Consistent resource attributes are how backends group data by service and environment.

Best Practices

Start Automatic, Then Add Business Context

Auto-instrumentation answers "which dependency is slow?". Manual spans and attributes answer "for which customers, plans, regions, and code paths?". Add attributes you'll filter on during incidents (tenant tier, feature flag, cache hit, payment provider), keeping them bounded.

Set service.name and Version Everywhere

Without service.name, data lands as unknown_service. Include service.version (or the git SHA) so you can correlate regressions with deploys.

Record Errors Properly

Call record_exception and set span status to ERROR for failures, but don't mark client-side 4xx as server errors in server spans. Follow the semantic conventions so error rates are computed consistently.

Keep Overhead in Check

Use the batch span processor (not simple or synchronous), sample appropriately (see sampling), avoid spans in tight loops, and don't put large payloads in attributes.

Common Mistakes

High-Cardinality Span Names and Metric Attributes

Forgetting to End Spans

Manually started spans (start_span without a context manager) that never call end() leak memory and never get exported. Prefer with tracer.start_as_current_span(...) or equivalent scoped APIs.

Not Flushing on Shutdown

Short-lived processes (CLI tools, serverless functions, batch jobs) exit before the batch processor exports. Call provider.shutdown() or force_flush() before exit. Serverless environments need special handling, such as the Lambda layer's extension.

FAQ

Should I use auto-instrumentation or manual instrumentation?

Both. Auto-instrumentation gives broad coverage of frameworks and dependencies with no code changes, and manual instrumentation adds business-level spans, attributes, and metrics that no agent can infer. Start with auto, then add manual where investigations need more context.

Does OpenTelemetry replace Prometheus client libraries?

It can. The OpenTelemetry metrics SDK exports to Prometheus via a scrape endpoint or OTLP (Prometheus 3.x ingests OTLP natively). Teams standardizing on OpenTelemetry for traces often adopt its metrics too, for one consistent API. Existing Prometheus instrumentation keeps working alongside it.

What's the performance overhead?

Usually low single-digit percentages of CPU and latency for typical web services with batching and sampling. Overhead grows with span volume, so very chatty instrumentation (spans per loop iteration, huge attributes) is the main risk. Measure in your environment.

How do I instrument a library I maintain?

Depend only on the OpenTelemetry API, not the SDK. Create a tracer and meter named after your library, and follow semantic conventions. Applications that configure an SDK will receive your telemetry, and those that don't pay essentially nothing.

Related Topics

References