Circuit Breakers & Resilience Patterns

In a distributed system, something is always failing: a database slows down, a dependency deploys a bug, a network link drops packets. Without protection, one slow dependency can exhaust threads and connections in every caller, which then time out their own callers, a cascading failure that takes down healthy services too. Resilience patterns contain failures: timeouts bound waiting, retries absorb transient errors, circuit breakers stop calling a dependency that's clearly failing, bulkheads isolate resources, and fallbacks degrade gracefully.

These patterns are essential in microservices, and useful anywhere code calls over a network. They're implemented in libraries (Resilience4j, Polly, failsafe-go, and Hystrix before it was retired) or in infrastructure (Envoy, service meshes).

TL;DR

Quick Example

Resilience4j in a Spring Boot service calling a recommendations API:

Core Concepts

Timeouts

Every network call needs a connect timeout and a request (read) timeout. Derive them from the caller's overall latency budget: if a request must finish in 1s and makes two sequential calls, neither can take 1s. Deadline propagation passes the remaining budget downstream (gRPC deadlines do this natively), so downstream services don't keep working on requests the caller has already abandoned.

Retries

Retries turn transient blips into successes, and they amplify load during real outages. Rules:

Circuit Breaker States

Benefits: callers stop wasting threads and time on a failing dependency, the dependency gets breathing room to recover, and users get fast responses (fallbacks or errors) rather than timeouts.

Bulkheads

Named after ship compartments: isolate resources per dependency, so one failing dependency can't sink the whole service:

Fallbacks and Graceful Degradation

When a dependency is unavailable, return something useful: cached data (possibly stale), defaults (generic recommendations), partial responses (the page without the reviews widget), or queuing work for later. Not everything can degrade (a payment can't succeed without the payment provider), but many features can, and users prefer a partially working page to an error page.

Load Shedding and Rate Limiting

When a service is overloaded, accepting more work makes everything slower until nothing succeeds. Load shedding rejects excess requests early (fast 503 or 429 responses), prioritizing critical traffic. Rate limiting caps per-client usage. Adaptive concurrency limits (for example Netflix's concurrency-limits, and Envoy's adaptive concurrency) adjust limits from observed latency.

Library vs Infrastructure

Many teams combine them: mesh-level timeouts, retries, and outlier detection, plus library-level fallbacks where business logic matters.

Best Practices

Set Timeouts Everywhere, Consistently

Audit every HTTP client, database driver, and messaging client for timeouts. Defaults are frequently infinite, or far too long.

Make Breakers Observable

Export circuit breaker state changes, failure rates, rejected calls, and retry counts as metrics, and alert when breakers open. An open breaker is a symptom worth investigating. See monitoring.

Test Failure Modes Deliberately

Inject latency, errors, and outages in staging and, carefully, in production (chaos engineering), and verify that timeouts fire, breakers open, fallbacks work, and alerts trigger.

Tune From Data

Base thresholds on observed latency percentiles and error rates. Breakers that open too eagerly cause self-inflicted outages; ones that never open provide no protection.

Common Mistakes

Retrying Non-Idempotent Operations

Retrying a timed-out "charge card" call without an idempotency key can charge customers twice. Make writes idempotent, or don't retry them automatically.

Retries Without Backoff and Jitter

Immediate retries from thousands of clients hit a recovering service at the same instant, knocking it over again. Always use backoff with jitter.

Circuit Breakers Without Timeouts

A breaker counts failures, but if calls hang for 60 seconds before failing, threads are exhausted long before the breaker opens. Timeouts come first.

FAQ

What does a circuit breaker do?

It monitors calls to a dependency, and when failures or slow responses exceed a threshold, it "opens", immediately failing further calls (or serving fallbacks) instead of waiting on a broken dependency. After a cooldown, it lets trial requests through, and closes again if they succeed. It prevents cascading failures and gives dependencies time to recover.

What's the difference between a retry and a circuit breaker?

Retries handle brief, transient failures by trying again. Circuit breakers handle sustained failures by stopping attempts altogether for a while. They complement each other: retries inside a closed circuit, and no retries once the circuit is open.

Where should resilience logic live: in code or in the service mesh?

Both have roles. Meshes provide consistent, language-agnostic timeouts, retries, and outlier detection. Application libraries add business-aware fallbacks and fine-grained control. Avoid duplicating retries across both layers, which multiplies load.

What is a bulkhead?

A pattern that partitions resources (threads, connections, concurrency slots) per dependency or workload, so a failure or slowdown in one area can't consume resources needed by others, like watertight compartments in a ship.

Related Topics

References