Redis Replication, Sentinel & Cluster

A single Redis instance is fast, but it's also a single point of failure and bounded by one machine's memory and CPU. Redis has three building blocks for going further. Replication copies data to replicas for read scaling and redundancy. Sentinel watches a primary and promotes a replica automatically when it fails. Redis Cluster shards data across multiple primaries, each with its own replicas, for both scale and availability.

Which you need depends on the problem. High availability with a dataset that fits on one node calls for replication plus Sentinel, or a managed service's failover. A dataset or write load bigger than one node calls for Cluster. Each adds constraints your application has to respect.

TL;DR

Quick Example

Creating a six-node cluster (three primaries, three replicas) and seeing slot routing:

Hash tags keep related keys on one shard so multi-key operations work:

Core Concepts

Replication

A replica connects to a primary, receives a full snapshot (RDB) on first sync, then streams every write command. If the link drops briefly, partial resynchronization resumes from a replication backlog instead of re-copying everything.

Sentinel

Sentinel is a separate process, usually run as three or more instances on different hosts, that:

  1. Monitors primaries and replicas with health checks.
  2. Agrees a primary is down when a quorum of Sentinels can't reach it.
  3. Fails over by promoting the best replica and reconfiguring the others to follow it.
  4. Serves discovery: clients ask Sentinel for the current primary's address.

Sentinel gives high availability to a non-sharded deployment. All data still lives on one primary.

Redis Cluster

Cluster shards data automatically:

Hash Tags

If a key contains {...}, only the part inside the braces is hashed. {order:991}:items and {order:991}:status land in the same slot, so transactions, Lua scripts, and multi-key commands work across them. Use tags deliberately: tagging too much onto one value creates a hot slot that defeats sharding.

Choosing a Topology

Redis is single-threaded for command execution (with I/O threads available), so even a large machine eventually CPU-bottlenecks on one core. Cluster scales writes by adding primaries. Valkey, the Linux Foundation fork of Redis, supports the same replication, Sentinel, and Cluster model.

Best Practices

Use Sentinel- or Cluster-Aware Clients

Hard-coding a primary's IP means an outage lasts until someone edits config. Use clients that discover the primary through Sentinel or follow cluster redirects and refresh topology (redis-py's RedisCluster, ioredis's Cluster, go-redis's ClusterClient, Lettuce).

Design Keys for Sharding From the Start

Decide which keys must be operated on together, and give them a shared hash tag. Retrofitting hash tags means renaming keys in production.

Spread Failure Domains

Place replicas in different availability zones from their primaries, and Sentinels on separate hosts. Keep an odd number of Sentinels (3 or 5) so quorum works. See high availability.

Test Failover

Trigger failovers deliberately (SENTINEL FAILOVER, CLUSTER FAILOVER on a replica, or managed-service test failover) and watch how your application behaves: reconnect times, error spikes, and lost writes. Chaos engineering applies here too.

Common Mistakes

CROSSSLOT Errors After Moving to Cluster

Audit MGET, MSET, transactions, Lua scripts, SUNION, and similar commands before migrating.

Big or Hot Keys in a Cluster

Sharding distributes keys, not the contents of one key. A single 5 GB sorted set or one extremely hot counter lives on one shard and can overload it. Split big collections and hot counters across several keys.

Assuming Replication Means No Data Loss

Asynchronous replication plus failover can drop the last writes. Split-brain situations can also accept writes on an isolated old primary; min-replicas-to-write limits this by refusing writes when too few replicas are connected.

FAQ

Do I need Sentinel if I use Redis Cluster?

No. Cluster has built-in failure detection and failover. Sentinel is for non-clustered primary-replica setups.

How many nodes does Redis Cluster need?

At least three primaries for a functioning cluster, and usually one replica each for availability, so six nodes is the common minimum for production. More primaries add capacity; more replicas per primary add redundancy and read capacity.

Can I read from replicas in Cluster?

Yes. Cluster-aware clients can route reads to replicas after sending READONLY. Accept that reads may be slightly stale, and keep writes and read-your-writes flows on primaries.

Should I run Redis Cluster myself or use a managed service?

Managed services take care of provisioning, patching, failover, backups, and resharding, which is usually worth it unless you have strong cost, control, or on-prem requirements. Self-hosting is very doable, but expect to own monitoring, upgrades, and failover testing.

Related Topics

References