Microservices Communication
Splitting a system into microservices turns in-process function calls into network calls, and networks are slow, unreliable, and asynchronous. How services communicate (synchronously over HTTP or gRPC, or asynchronously through messages and events) shapes latency, availability, coupling, and how failures spread. Many microservice failures, such as cascading outages, distributed monoliths, and inconsistent data, trace back to communication design.
The core decision is synchronous vs asynchronous. Synchronous calls are simple and give immediate answers, but they couple availability: if B is down, A fails. Asynchronous messaging decouples services in time, and absorbs load spikes, at the cost of eventual consistency and harder debugging. Most real systems use both, deliberately.
TL;DR
- Synchronous: REST/HTTP (universal, simple) or gRPC (typed, fast, streaming), best for queries needing an immediate answer.
- Asynchronous: commands to queues and events to topics (Kafka, RabbitMQ, SQS/SNS, NATS), best for workflows, notifications, and decoupling.
- Every synchronous call needs timeouts, bounded retries with backoff and jitter, and idempotency; add circuit breakers to stop cascades.
- Put an API gateway (and optionally BFFs) at the edge; keep internal traffic service-to-service or via a mesh.
- Avoid long synchronous call chains: each hop multiplies latency and failure probability.
- Manage contracts explicitly (OpenAPI, Protobuf, AsyncAPI, schema registries) with backward-compatible evolution and contract tests.
Quick Example
Checkout combining sync and async communication:
The user waits only for pricing and payment. Stock reservation, email, analytics, and loyalty happen asynchronously and independently, so a slow email service never delays checkout.
Core Concepts
Synchronous Communication
Synchronous calls couple availability and latency: the caller's response time includes every downstream call, and the caller's availability is roughly the product of its dependencies' availability. Five services at 99.9% each chained together gives about 99.5%.
Asynchronous Communication
- Commands ("ReserveStock"): sent to a specific service via a queue, expecting it to act.
- Events ("OrderPlaced"): facts published to a topic, with zero or more subscribers reacting independently. The publisher doesn't know or care who consumes them.
- Request-reply over messaging: a reply queue plus a correlation ID, for async interactions needing responses.
Benefits: temporal decoupling (consumers can be down temporarily), load leveling (queues absorb spikes), and easy fan-out to new consumers. Costs: eventual consistency, duplicate and out-of-order delivery, harder tracing, and operating a broker. See message queues, event-driven architecture, and Kafka.
Choosing Sync vs Async
Resilience for Network Calls
- Timeouts on every call, set per dependency from latency budgets. Never rely on library defaults, which are often infinite.
- Retries only for transient errors and idempotent operations, with exponential backoff, jitter, and caps. Retry budgets stop retry storms.
- Idempotency keys so retried writes don't duplicate effects. See idempotency.
- Circuit breakers and bulkheads to fail fast and isolate failing dependencies. See circuit breakers.
- Fallbacks and graceful degradation: cached data, defaults, or partial responses.
Edge Patterns: API Gateway and BFF
- An API gateway is the single entry point for clients: routing, TLS, authentication, rate limiting, request aggregation, and observability.
- Backends-for-frontends (BFFs) are gateway-like services tailored to a specific client (web, mobile), aggregating calls and shaping responses. They reduce chattiness over slow mobile networks.
Internally, service meshes can add mTLS, retries, and traffic policies without library code. See service mesh.
Contracts and Versioning
- Define contracts explicitly: OpenAPI for REST, Protobuf for gRPC, AsyncAPI and schema registries (Avro, Protobuf, JSON Schema) for events.
- Evolve backward-compatibly: add optional fields; never remove or repurpose fields consumers use. Version when breaking changes are unavoidable. See API versioning.
- Consumer-driven contract tests (Pact) verify that providers don't break consumers, without full end-to-end environments.
Best Practices
Minimize Synchronous Call Depth
Design so a user request touches few services synchronously. Deep chains (A→B→C→D) multiply latency and failure rates. Denormalize data via events, or restructure service boundaries.
Prefer Events for Cross-Domain Side Effects
When an action in one domain triggers work in others, publish an event rather than calling each service. New consumers can then be added without changing the publisher.
Propagate Context
Pass trace context, correlation IDs, and deadlines across calls and messages, so failures can be traced end to end. See distributed tracing.
Design for Duplicates and Reordering
Asynchronous consumers must handle messages delivered more than once and occasionally out of order: idempotent handlers, version checks, and upserts.
Common Mistakes
The Distributed Monolith
Services that must all be deployed together, call each other synchronously for every operation, and share databases have microservice costs without the benefits. Revisit boundaries, and reduce synchronous coupling.
No Timeouts
A downstream service hanging with no caller timeouts ties up threads and connections everywhere, and one slow dependency takes down the whole system.
Chatty Interfaces
Fetching a list, then calling another service once per item (the N+1 problem across the network) destroys performance. Provide batch endpoints, or aggregate data via events.
FAQ
Should microservices communicate with REST or gRPC?
REST is universal, simple to debug, and great for public and edge APIs. gRPC offers strongly typed contracts, better performance, and streaming, which suits internal service-to-service traffic. Many organizations use REST or GraphQL at the edge and gRPC internally.
When should microservices use messaging instead of HTTP?
When the caller doesn't need an immediate result, several services react to the same occurrence, workloads are bursty, or you want services to keep working when others are temporarily unavailable. Use HTTP or gRPC for queries and interactions that need an answer now.
What is an API gateway for?
It's the single entry point for external clients, handling cross-cutting concerns (routing, authentication, rate limiting, TLS, aggregation, logging) so individual services don't each implement them. It also shields clients from internal service topology changes.
How do I avoid cascading failures between services?
Set timeouts on every call, use circuit breakers and bulkheads, limit retries with backoff and jitter, provide fallbacks, prefer asynchronous communication for non-critical work, and load-test failure scenarios. Chaos engineering helps validate these defenses.
Related Topics
- Microservices — The architecture overview
- Circuit Breakers — Resilience for synchronous calls
- Service Discovery — Finding services to call
- Event-Driven Architecture — Asynchronous integration
- gRPC — High-performance RPC between services
- API Gateway — The edge entry point