Service Discovery
In a static world, service A calls service B at a fixed hostname and port. In microservices running on autoscaling infrastructure, instances come and go constantly: deployments replace them, autoscalers add and remove them, failures kill them, and IP addresses change every time. Service discovery answers "where are the healthy instances of B right now?", continuously and automatically.
On Kubernetes, discovery is largely built in through Services and DNS. Outside Kubernetes, or across clusters and hybrid environments, registries like Consul, Eureka, and cloud-native discovery services fill the role. Service meshes build on discovery to add load balancing, mTLS, and traffic policy.
TL;DR
- A service registry tracks available instances and their locations; clients query it (directly or through a proxy).
- Client-side discovery: the client queries the registry and load-balances itself. Server-side discovery: the client calls a load balancer or proxy that does the lookup.
- Instances register via self-registration (the app registers itself) or third-party registration (the platform registers it).
- Health checks remove unhealthy instances from the registry.
- Kubernetes: Services, EndpointSlices, and cluster DNS provide discovery out of the box.
- Service meshes (Istio, Linkerd, Consul) layer smart load balancing, retries, and mTLS on discovery; multi-cluster setups need federation.
Quick Example
Discovery on Kubernetes: a Service gives a stable name for a changing set of pods.
Registration with Consul outside Kubernetes (agent service definition):
Core Concepts
The Service Registry
A registry is a database of service instances: name, address, port, metadata (version, zone, tags), and health. It must be highly available, since it's on the critical path of every call. That's why registries like Consul, etcd, and ZooKeeper use consensus, and why clients cache results.
Client-Side vs Server-Side Discovery
Service meshes blur the line: a sidecar proxy next to each service does client-side-style balancing, transparently to the application.
Registration Patterns
- Self-registration: the service registers on startup and deregisters on shutdown, sending heartbeats (Eureka-style). It's simple, but couples the app to the registry.
- Third-party registration: the platform registers instances automatically (Kubernetes, ECS with Cloud Map, Consul catalog sync). Apps stay unaware of discovery. This is usually preferable.
Health Checking
Registries must stop routing to unhealthy instances:
- Active checks: the registry or agent probes HTTP, TCP, or gRPC health endpoints.
- Heartbeats / TTLs: instances renew their registration periodically, and silence means removal.
- Readiness vs liveness: route only to instances ready to serve. In Kubernetes, readiness probes control Service endpoints. See Kubernetes workloads.
Deregister promptly on shutdown, and drain connections, to avoid errors during deployments.
DNS-Based Discovery
DNS is the universal discovery interface: A or AAAA records for instance IPs, SRV records for ports, and short TTLs for freshness. Kubernetes (CoreDNS), Consul, and AWS Cloud Map all expose DNS. Limitations: client DNS caching can serve stale addresses, and plain DNS doesn't carry health or metadata richly. Many clients combine DNS with health-aware registries or proxies. See DNS.
Kubernetes Discovery
- Services provide stable virtual IPs and DNS names (
svc.namespace.svc.cluster.local). - EndpointSlices list ready pod IPs behind each Service, updated as pods change.
- Headless Services return pod IPs directly, for client-side balancing (gRPC) or StatefulSets.
- Cross-namespace calls use fully qualified names, and NetworkPolicies control who may call whom.
Service Mesh and Multi-Cluster
A service mesh consumes discovery data to provide per-request load balancing (least requests, locality-aware), retries and timeouts, mTLS identity, and traffic splitting for canaries. For multi-cluster or hybrid deployments, discovery must span clusters: mesh multi-cluster modes, Consul federation, Kubernetes multi-cluster Services (MCS API), or global load balancers with health-aware DNS.
Best Practices
Let the Platform Handle Registration
On Kubernetes, rely on Services and readiness probes rather than embedding registry clients in apps. Outside Kubernetes, use agents or orchestrator integrations (Consul with Nomad, ECS with Cloud Map) for third-party registration.
Make Health Checks Meaningful
Readiness should reflect ability to serve (warm-up done, critical dependencies reachable), and liveness only process health. Poor checks route traffic to broken instances, or remove healthy ones during dependency blips.
Cache and Degrade Gracefully
Clients should cache discovery results and keep using last-known-good instances if the registry is briefly unavailable. A registry outage shouldn't become a total outage.
Prefer Locality-Aware Routing
Route to instances in the same zone when possible, which reduces latency and cross-zone data transfer costs, with failover to other zones when local capacity is unhealthy.
Common Mistakes
Hard-Coding IPs or Hostnames per Instance
Static lists of instance addresses break with every deploy and autoscaling event. Use service names resolved by discovery.
Long DNS TTLs and Client Caching
Clients (notably older JVMs with default DNS caching) holding stale addresses keep calling terminated instances. Use short TTLs, and configure client DNS caching appropriately.
Long-Lived Connections Ignoring New Instances
HTTP/2 and gRPC connections established to a few instances don't rebalance when new ones appear. Use client-side load balancing with periodic re-resolution, connection max-age settings, or a mesh.
FAQ
What is service discovery?
The mechanism by which services locate the network addresses of healthy instances of other services in dynamic environments where instances start, stop, and move frequently. It typically involves a registry of instances, health checking, and a lookup method (DNS, API, or a proxy).
Do I need Consul or Eureka on Kubernetes?
Usually not. Kubernetes Services, EndpointSlices, and cluster DNS provide discovery and basic load balancing out of the box. Consul or a mesh becomes useful for multi-cluster, hybrid (VMs plus Kubernetes), or advanced traffic management scenarios.
What's the difference between client-side and server-side discovery?
In client-side discovery, the calling service queries the registry and chooses an instance itself. In server-side discovery, the caller sends requests to a load balancer or proxy that looks up instances and routes the request. Server-side keeps clients simple; client-side avoids an extra hop and enables smarter balancing.
How do health checks relate to discovery?
Discovery should only return instances able to serve traffic. Health checks (active probes, or heartbeats) mark instances healthy or unhealthy, and registries and load balancers exclude unhealthy ones automatically, so failures don't reach callers.
Related Topics
- Microservices — The architecture overview
- Microservices Communication — Calling discovered services
- Kubernetes Networking — Services, DNS, and EndpointSlices
- Service Mesh — Discovery plus traffic management
- Load Balancing — Distributing requests across instances
- DNS — The universal discovery protocol