Redis Persistence

Redis keeps its entire dataset in memory. That's why it's fast, and why a restart could wipe everything. Persistence writes data to disk so Redis can reload it after a restart or crash. Redis offers two mechanisms, RDB snapshots and the append-only file (AOF), which can be used alone, together, or not at all.

The right setting depends on what Redis holds. A pure cache can often skip persistence entirely. Sessions, queues, rate-limit state, or data where Redis is the primary store need durability, and you have to choose how much data you can afford to lose in a crash versus how much write performance you'll trade for it.

TL;DR

Quick Example

A durable configuration for Redis as a primary store for queues and sessions:

Core Concepts

RDB Snapshots

Redis forks a child process that writes the entire dataset to a compact binary file, while the parent keeps serving requests. The OS's copy-on-write means memory pages are duplicated only as the parent modifies them during the snapshot.

Append-Only File (AOF)

Every write command is appended to a log. On restart, Redis replays the log to rebuild the dataset. Durability depends on the fsync policy:

AOF Rewrite and Multi-Part AOF

The log grows forever, since INCR counter a million times is a million entries. AOF rewrite (BGREWRITEAOF, or automatic by size thresholds) compacts it into the minimal commands needed to recreate the current state. Since Redis 7, AOF is multi-part: a base file (RDB or AOF format) plus incremental files, tracked by a manifest in an appendonlydir/ directory.

Hybrid Persistence

With aof-use-rdb-preamble yes (the default), rewrites produce a base file in RDB format followed by AOF increments. Restarts load the compact snapshot quickly, then replay only the recent commands.

Choosing a Strategy

Managed services (ElastiCache, Azure Cache for Redis, Redis Cloud, Upstash) expose these as snapshot and AOF settings, and add automated backups.

Backups and Restores

Best Practices

Leave Memory Headroom for Forks

On write-heavy workloads, copy-on-write during a snapshot or rewrite can nearly double memory use in the worst case. Set maxmemory well below physical RAM (for example 60–75%), and set vm.overcommit_memory = 1 on Linux so forks don't fail.

Disable Transparent Huge Pages

THP makes copy-on-write copy 2 MB pages instead of 4 KB ones, inflating memory and latency during forks. Redis warns about it at startup; disable it on Redis hosts.

Monitor Persistence Health

Alert on rdb_last_bgsave_status:err, aof_last_write_status:err, growing aof_delayed_fsync, and long fork times (latest_fork_usec). A failed snapshot can make Redis refuse writes (stop-writes-on-bgsave-error yes), which surfaces as sudden application errors.

Combine Persistence With Replication

Replicas provide availability and fast failover (via Sentinel or Cluster); persistence provides recovery after full restarts. You usually want both for important data, but don't mistake one for the other.

Common Mistakes

Assuming a Cache Instance Is Durable

Teams start storing sessions or job queues in "the cache" Redis, which has persistence off and an eviction policy on. A restart or memory pressure then silently drops user sessions or jobs. Separate durable data onto an instance configured for it.

Replica Without Persistence on a Primary Without Persistence

If a primary with persistence disabled restarts automatically, it comes back empty, and replicas then sync that empty dataset, wiping their copies too. Either enable persistence on the primary or disable its automatic restart.

Relying on SAVE or Frequent Snapshots for Durability

Snapshotting every few seconds on a large dataset means constant forking, CPU load, and latency spikes. Use AOF for fine-grained durability and RDB at sensible intervals for backups.

FAQ

RDB or AOF: which should I use?

For data you care about, use both: AOF (everysec) for durability and RDB snapshots for backups and fast restarts. For a rebuildable cache, RDB alone, or no persistence at all, is fine. AOF alone without backups leaves you without easy point-in-time copies.

How much data can I lose with appendfsync everysec?

Typically up to about one second of writes if the machine crashes or loses power. If the disk falls behind, Redis can briefly delay fsyncs, and the window can grow to around two seconds. A clean shutdown loses nothing.

Does persistence slow Redis down?

Somewhat. everysec AOF adds minor overhead for most workloads, while always significantly reduces write throughput. Forking for RDB or rewrites can cause latency spikes on large datasets. Fast disks, memory headroom, and scheduling heavy snapshots during off-peak hours all help.

Is Redis a real database or just a cache?

It can be either. With AOF, replication, and backups, Redis serves as a primary store for many workloads, but its data model and memory-bound size differ from relational databases. Many teams keep the system of record in PostgreSQL and use Redis for fast, derived, or ephemeral state.

Related Topics

References