Databases 20 min read

Illustrated Redis: Architecture, Replication & Persistence Internals

This article provides a comprehensive visual guide to Redis architecture, covering single-instance deployment, high-availability setups with Sentinel, Redis Cluster sharding with hash slots, replication mechanics using replication IDs and offsets, and persistence models including RDB snapshots, AOF logging, and the fork/copy-on-write mechanism.

Architect's Guide
Architect's Guide
Architect's Guide
Illustrated Redis: Architecture, Replication & Persistence Internals

What Is Redis?

Redis (REmote DIctionary Service) is an open-source key-value database server. More accurately, it is a data structure server that organizes data by structure rather than iteration or sorting. While early usage resembled Memcached, Redis has evolved to support pub/sub, streams, and queues.

Primarily an in-memory database, Redis often sits in front of a primary database like MySQL or PostgreSQL to boost performance by leveraging fast memory access. It caches infrequently changing but frequently requested data, as well as less critical but frequently changing data such as session caches, leaderboards, and dashboard analytics.

With plugins and high-availability (HA) configurations, Redis can also serve as a mature primary database for certain workloads. It blurs the line between cache and data store, offering much faster read/write speeds than traditional disk-based databases.

Redis overview
Redis overview
Redis data structures
Redis data structures
Redis as cache
Redis as cache
Redis vs Memcached
Redis vs Memcached

Redis Architecture

Before diving into internals, the article outlines four main deployment topologies and their trade-offs:

Single Redis Instance

Redis High Availability (Master-Replica)

Redis Sentinel

Redis Cluster

Single Redis Instance

Single Redis instance
Single Redis instance

The simplest deployment. It enables quick setup and significant performance gains for caching with minimal configuration. However, if the instance fails, all client calls fail, reducing overall system availability.

Commands are processed in memory. If persistence is enabled, a forked child process periodically creates RDB (point-in-time snapshot) or AOF (append-only file) files. Without persistence, data is lost on restart. With persistence enabled, data is reloaded from RDB or AOF on startup.

Redis High Availability (Master-Replica)

Master-replica replication
Master-replica replication

Replicas stay synchronized with the master. Writes to the master are copied to replica output buffers. Replicas can scale reads and provide failover if the master is lost. This introduces distributed-system complexity.

Redis Replication

Each master has a replication ID and an offset. The offset increments with every operation. A replica only a few offsets behind receives missing commands and replays them (partial sync). If replication IDs disagree or the master doesn't know the offset, the replica requests a full sync: the master creates a new RDB snapshot, sends it, and buffers intermediate writes to send afterward.

For example, two instances (master and replica) share the same replication ID but differ by a few hundred commands. Replaying those commands makes their datasets identical. If replication IDs are completely different with no common ancestor, an expensive full sync is required. Knowing the previous replication ID allows inference of a common ancestor, making partial sync possible again.

Redis Sentinel

Redis Sentinel architecture
Redis Sentinel architecture

Sentinel is a distributed system where a group of Sentinel processes coordinate to provide HA for Redis. It avoids a single point of failure in the monitoring layer itself.

Responsibilities:

Monitoring — Ensure master and replicas are healthy.

Notification — Alert administrators of Redis events.

Failover Management — If the master is unavailable and a quorum of Sentinels agree, initiate failover.

Configuration Provider — Act as a discovery service for the current master.

Failover detection uses a quorum protocol among multiple Sentinels, increasing robustness. The article recommends running at least three Sentinel nodes with a quorum of two, and placing a Sentinel next to each application server to avoid network reachability issues.

Sentinel quorum and failure tolerance
Sentinel quorum and failure tolerance

Potential issues:

What if Sentinels exceed quorum?

Network partition isolating the old master in the minority — writes to that master are lost when the system recovers.

Misaligned network topology between Sentinels and application nodes.

No strong durability guarantees exist because replication is asynchronous. Data loss can occur when clients discover a new primary. Mitigation: configure the master to stop accepting writes if at least one replica hasn't acknowledged writes (using min-replicas-to-write and min-replicas-max-lag).

Redis Cluster

Redis Cluster topology
Redis Cluster topology

When data exceeds a single machine's memory (max 24 TiB on AWS), horizontal scaling is needed. Redis Cluster shards data across multiple nodes.

Each node holds a shard. To locate a key's shard, Redis Cluster uses algorithmic sharding: hash the key, modulo total shards. However, adding a new shard would require massive data movement. Instead, Redis Cluster uses hash slots (16,384 slots). Keys map to slots; slots map to nodes. Resharding moves slots between nodes, not individual keys, enabling zero-downtime scaling with minimal performance impact.

Example: Initially M1 holds slots 0–8191, M2 holds 8192–16383. Key "foo" hashes to a slot in M2. After adding M3, slots are redistributed: M1 0–5460, M2 5461–10922, M3 10923–16383. Only keys in moved slots are migrated; slot-to-key mapping remains stable.

Gossip Protocol

Cluster nodes continuously gossip to determine health. If enough nodes agree a master (e.g., M1) is down, its replica (S1) can be promoted. The required agreement count is configurable. To avoid split-brain, an odd number of masters with at least two replicas each is recommended for the most robust setup.

Redis Persistence Models

Understanding persistence is crucial when data safety matters. For caching or real-time analytics, occasional loss may be acceptable; for other cases, durability guarantees are needed.

Persistence models comparison
Persistence models comparison

No Persistence

Persistence can be fully disabled for maximum speed, with no durability guarantees.

RDB (Redis Database)

RDB performs point-in-time snapshots at configured intervals. Drawback: data between snapshots can be lost. Forking the main process for large datasets may cause brief latency spikes. However, RDB files load into memory much faster than AOF.

AOF (Append Only File)

AOF logs every write operation, replaying them on restart to reconstruct the dataset. Operations are buffered and periodically fsynced to disk (configurable). This provides better durability than RDB because it's append-only. Downsides: less compact format, larger disk usage.

RDB + AOF Combined

Both can be enabled simultaneously, trading speed for durability. On restart, Redis uses AOF to rebuild data because it's more complete.

Forking and Copy-on-Write

Fork and copy-on-write
Fork and copy-on-write

Redis leverages OS forking and copy-on-write (COW) to snapshot efficiently in a single-threaded process. Fork creates a child process sharing the parent's memory pages. The child performs the snapshot (RDB or AOF rewrite). If no writes occur during the snapshot, no new memory is allocated. When writes happen, the kernel copies only modified pages to new locations, leaving the child with a consistent snapshot. This allows gigabytes of memory to be snapshotted quickly with minimal extra memory overhead.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

ArchitectureHigh AvailabilityRedisPersistenceReplicationSentinelClusterCopy-on-Write
Architect's Guide
Written by

Architect's Guide

Dedicated to sharing programmer-architect skills—Java backend, system, microservice, and distributed architectures—to help you become a senior architect.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.