Illustrated Redis: Architecture, Replication & Persistence Internals
This article provides a comprehensive visual guide to Redis architecture, covering single-instance deployment, high-availability setups with Sentinel, Redis Cluster sharding with hash slots, replication mechanics using replication IDs and offsets, and persistence models including RDB snapshots, AOF logging, and the fork/copy-on-write mechanism.
What Is Redis?
Redis (REmote DIctionary Service) is an open-source key-value database server. More accurately, it is a data structure server that organizes data by structure rather than iteration or sorting. While early usage resembled Memcached, Redis has evolved to support pub/sub, streams, and queues.
Primarily an in-memory database, Redis often sits in front of a primary database like MySQL or PostgreSQL to boost performance by leveraging fast memory access. It caches infrequently changing but frequently requested data, as well as less critical but frequently changing data such as session caches, leaderboards, and dashboard analytics.
With plugins and high-availability (HA) configurations, Redis can also serve as a mature primary database for certain workloads. It blurs the line between cache and data store, offering much faster read/write speeds than traditional disk-based databases.
Redis Architecture
Before diving into internals, the article outlines four main deployment topologies and their trade-offs:
Single Redis Instance
Redis High Availability (Master-Replica)
Redis Sentinel
Redis Cluster
Single Redis Instance
The simplest deployment. It enables quick setup and significant performance gains for caching with minimal configuration. However, if the instance fails, all client calls fail, reducing overall system availability.
Commands are processed in memory. If persistence is enabled, a forked child process periodically creates RDB (point-in-time snapshot) or AOF (append-only file) files. Without persistence, data is lost on restart. With persistence enabled, data is reloaded from RDB or AOF on startup.
Redis High Availability (Master-Replica)
Replicas stay synchronized with the master. Writes to the master are copied to replica output buffers. Replicas can scale reads and provide failover if the master is lost. This introduces distributed-system complexity.
Redis Replication
Each master has a replication ID and an offset. The offset increments with every operation. A replica only a few offsets behind receives missing commands and replays them (partial sync). If replication IDs disagree or the master doesn't know the offset, the replica requests a full sync: the master creates a new RDB snapshot, sends it, and buffers intermediate writes to send afterward.
For example, two instances (master and replica) share the same replication ID but differ by a few hundred commands. Replaying those commands makes their datasets identical. If replication IDs are completely different with no common ancestor, an expensive full sync is required. Knowing the previous replication ID allows inference of a common ancestor, making partial sync possible again.
Redis Sentinel
Sentinel is a distributed system where a group of Sentinel processes coordinate to provide HA for Redis. It avoids a single point of failure in the monitoring layer itself.
Responsibilities:
Monitoring — Ensure master and replicas are healthy.
Notification — Alert administrators of Redis events.
Failover Management — If the master is unavailable and a quorum of Sentinels agree, initiate failover.
Configuration Provider — Act as a discovery service for the current master.
Failover detection uses a quorum protocol among multiple Sentinels, increasing robustness. The article recommends running at least three Sentinel nodes with a quorum of two, and placing a Sentinel next to each application server to avoid network reachability issues.
Potential issues:
What if Sentinels exceed quorum?
Network partition isolating the old master in the minority — writes to that master are lost when the system recovers.
Misaligned network topology between Sentinels and application nodes.
No strong durability guarantees exist because replication is asynchronous. Data loss can occur when clients discover a new primary. Mitigation: configure the master to stop accepting writes if at least one replica hasn't acknowledged writes (using min-replicas-to-write and min-replicas-max-lag).
Redis Cluster
When data exceeds a single machine's memory (max 24 TiB on AWS), horizontal scaling is needed. Redis Cluster shards data across multiple nodes.
Each node holds a shard. To locate a key's shard, Redis Cluster uses algorithmic sharding: hash the key, modulo total shards. However, adding a new shard would require massive data movement. Instead, Redis Cluster uses hash slots (16,384 slots). Keys map to slots; slots map to nodes. Resharding moves slots between nodes, not individual keys, enabling zero-downtime scaling with minimal performance impact.
Example: Initially M1 holds slots 0–8191, M2 holds 8192–16383. Key "foo" hashes to a slot in M2. After adding M3, slots are redistributed: M1 0–5460, M2 5461–10922, M3 10923–16383. Only keys in moved slots are migrated; slot-to-key mapping remains stable.
Gossip Protocol
Cluster nodes continuously gossip to determine health. If enough nodes agree a master (e.g., M1) is down, its replica (S1) can be promoted. The required agreement count is configurable. To avoid split-brain, an odd number of masters with at least two replicas each is recommended for the most robust setup.
Redis Persistence Models
Understanding persistence is crucial when data safety matters. For caching or real-time analytics, occasional loss may be acceptable; for other cases, durability guarantees are needed.
No Persistence
Persistence can be fully disabled for maximum speed, with no durability guarantees.
RDB (Redis Database)
RDB performs point-in-time snapshots at configured intervals. Drawback: data between snapshots can be lost. Forking the main process for large datasets may cause brief latency spikes. However, RDB files load into memory much faster than AOF.
AOF (Append Only File)
AOF logs every write operation, replaying them on restart to reconstruct the dataset. Operations are buffered and periodically fsynced to disk (configurable). This provides better durability than RDB because it's append-only. Downsides: less compact format, larger disk usage.
RDB + AOF Combined
Both can be enabled simultaneously, trading speed for durability. On restart, Redis uses AOF to rebuild data because it's more complete.
Forking and Copy-on-Write
Redis leverages OS forking and copy-on-write (COW) to snapshot efficiently in a single-threaded process. Fork creates a child process sharing the parent's memory pages. The child performs the snapshot (RDB or AOF rewrite). If no writes occur during the snapshot, no new memory is allocated. When writes happen, the kernel copies only modified pages to new locations, leaving the child with a consistent snapshot. This allows gigabytes of memory to be snapshotted quickly with minimal extra memory overhead.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Architect's Guide
Dedicated to sharing programmer-architect skills—Java backend, system, microservice, and distributed architectures—to help you become a senior architect.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
