Redis Replication, Sentinel & Cluster
A single Redis instance is fast, but it's also a single point of failure and bounded by one machine's memory and CPU. Redis has three building blocks for going further. Replication copies data to replicas for read scaling and redundancy. Sentinel watches a primary and promotes a replica automatically when it fails. Redis Cluster shards data across multiple primaries, each with its own replicas, for both scale and availability.
Which you need depends on the problem. High availability with a dataset that fits on one node calls for replication plus Sentinel, or a managed service's failover. A dataset or write load bigger than one node calls for Cluster. Each adds constraints your application has to respect.
TL;DR
- Replication is asynchronous: replicas copy the primary's writes. Acknowledged writes can be lost on failover.
- Sentinel monitors primaries, agrees on failure by quorum, promotes a replica, and tells clients the new primary.
- Redis Cluster splits the keyspace into 16,384 hash slots distributed across primaries; each primary has replicas and failover is built in.
- Multi-key commands in Cluster only work when all keys map to the same slot; use hash tags like
{user:42}. - Clients must be Sentinel-aware or cluster-aware to follow failovers and slot redirects (
MOVED/ASK). - Managed services (ElastiCache, MemoryDB, Azure Managed Redis, Redis Cloud) handle most of this operationally.
Quick Example
Creating a six-node cluster (three primaries, three replicas) and seeing slot routing:
Hash tags keep related keys on one shard so multi-key operations work:
Core Concepts
Replication
A replica connects to a primary, receives a full snapshot (RDB) on first sync, then streams every write command. If the link drops briefly, partial resynchronization resumes from a replication backlog instead of re-copying everything.
- Replication is asynchronous. The primary acknowledges writes before replicas have them, so a failover can lose the most recent writes.
WAIT numreplicas timeoutblocks until N replicas have acknowledged, which reduces but doesn't eliminate the risk. - Replicas are read-only by default and can serve reads, with the possibility of slightly stale data.
- See database replication for the general trade-offs.
Sentinel
Sentinel is a separate process, usually run as three or more instances on different hosts, that:
- Monitors primaries and replicas with health checks.
- Agrees a primary is down when a quorum of Sentinels can't reach it.
- Fails over by promoting the best replica and reconfiguring the others to follow it.
- Serves discovery: clients ask Sentinel for the current primary's address.
Sentinel gives high availability to a non-sharded deployment. All data still lives on one primary.
Redis Cluster
Cluster shards data automatically:
- The keyspace is divided into 16,384 hash slots;
slot = CRC16(key) mod 16384. - Each primary owns a range of slots, and each has replicas that take over if it fails. Nodes gossip on a cluster bus and vote on failures, so no Sentinel is needed.
- A client sending a command to the wrong node gets a
MOVEDredirect with the right node. Cluster-aware clients cache the slot map and route directly. - Resharding moves slots between nodes online (
redis-cli --cluster reshard/rebalance), usingASKredirects while a slot is mid-migration.
Hash Tags
If a key contains {...}, only the part inside the braces is hashed. {order:991}:items and {order:991}:status land in the same slot, so transactions, Lua scripts, and multi-key commands work across them. Use tags deliberately: tagging too much onto one value creates a hot slot that defeats sharding.
Choosing a Topology
Redis is single-threaded for command execution (with I/O threads available), so even a large machine eventually CPU-bottlenecks on one core. Cluster scales writes by adding primaries. Valkey, the Linux Foundation fork of Redis, supports the same replication, Sentinel, and Cluster model.
Best Practices
Use Sentinel- or Cluster-Aware Clients
Hard-coding a primary's IP means an outage lasts until someone edits config. Use clients that discover the primary through Sentinel or follow cluster redirects and refresh topology (redis-py's RedisCluster, ioredis's Cluster, go-redis's ClusterClient, Lettuce).
Design Keys for Sharding From the Start
Decide which keys must be operated on together, and give them a shared hash tag. Retrofitting hash tags means renaming keys in production.
Spread Failure Domains
Place replicas in different availability zones from their primaries, and Sentinels on separate hosts. Keep an odd number of Sentinels (3 or 5) so quorum works. See high availability.
Test Failover
Trigger failovers deliberately (SENTINEL FAILOVER, CLUSTER FAILOVER on a replica, or managed-service test failover) and watch how your application behaves: reconnect times, error spikes, and lost writes. Chaos engineering applies here too.
Common Mistakes
CROSSSLOT Errors After Moving to Cluster
Audit MGET, MSET, transactions, Lua scripts, SUNION, and similar commands before migrating.
Big or Hot Keys in a Cluster
Sharding distributes keys, not the contents of one key. A single 5 GB sorted set or one extremely hot counter lives on one shard and can overload it. Split big collections and hot counters across several keys.
Assuming Replication Means No Data Loss
Asynchronous replication plus failover can drop the last writes. Split-brain situations can also accept writes on an isolated old primary; min-replicas-to-write limits this by refusing writes when too few replicas are connected.
FAQ
Do I need Sentinel if I use Redis Cluster?
No. Cluster has built-in failure detection and failover. Sentinel is for non-clustered primary-replica setups.
How many nodes does Redis Cluster need?
At least three primaries for a functioning cluster, and usually one replica each for availability, so six nodes is the common minimum for production. More primaries add capacity; more replicas per primary add redundancy and read capacity.
Can I read from replicas in Cluster?
Yes. Cluster-aware clients can route reads to replicas after sending READONLY. Accept that reads may be slightly stale, and keep writes and read-your-writes flows on primaries.
Should I run Redis Cluster myself or use a managed service?
Managed services take care of provisioning, patching, failover, backups, and resharding, which is usually worth it unless you have strong cost, control, or on-prem requirements. Self-hosting is very doable, but expect to own monitoring, upgrades, and failover testing.
Related Topics
- Redis — The database overview
- Redis Persistence — Durability alongside replication
- Database Sharding — Sharding concepts in general
- Database Replication — Synchronous vs asynchronous trade-offs
- High Availability — Designing for failover
- AWS ElastiCache — Managed Redis/Valkey on AWS