Redis Distributed Locks: 7 Critical Pitfalls and How to Avoid Them
This article details seven common pitfalls in Redis distributed locking, covering atomic acquisition, ownership verification, lease expiration, watchdog limitations, replication risks, Redlock assumptions, and the need for fencing tokens, with code examples and comparisons to ZooKeeper and etcd.
Redis distributed locks are widely used but often implemented incorrectly. This article, based on 11 years of e-commerce and payment backend experience, identifies seven pitfalls that cause either availability issues (efficiency locks) or data corruption (correctness locks). It references Martin Kleppmann's 2016 paper "How to do distributed locking" and includes his original diagrams (CC BY 3.0).
Consensus: Efficiency vs Correctness Locks
Redis locks solve two distinct problem classes:
Efficiency locks : Prevent duplicate work (e.g., sending daily emails). Single-node Redis with Redisson is sufficient.
Correctness locks : Enforce mutual exclusion on shared resources (e.g., inventory deduction). Redis locks can only assist; the real guarantee must come from the resource side via optimistic locking ( UPDATE ... WHERE version = $old).
Use Redis locks for deduplication; for strict mutual exclusion, combine with resource-side version control.
Pitfall 1: SETNX + EXPIRE Is Two Steps — Crash Between Them Causes Deadlock
Classic buggy code:
jedis.setnx("lock:order:42", "1");
jedis.expire("lock:order:42", 30);If the client crashes after setnx succeeds but before expire runs, the key remains forever, blocking all future acquisitions. Since Redis 2.6.12 (2013), the atomic SET key <uuid> NX PX 30000 combines both operations. NX ensures mutual exclusion, PX sets TTL in milliseconds, and <uuid> is a unique ownership token (20-byte random from /dev/urandom, or UUID v4 / clientId+threadId). Never use predictable values like "1", IP, or timestamps.
Treat setnx as deprecated; SET ... NX ... PX replaces it entirely.
Pitfall 2: Deleting Without UUID Verification Deletes Others' Locks
Releasing with a plain DEL is unsafe because the lock may have expired and been reacquired by another client. The correct release uses a Lua script (or Redis 8.4's DELEX key IFEQ my_random_value) that atomically checks ownership before deletion:
if redis.call("get", KEYS[1]) == ARGV[1] then
return redis.call("del", KEYS[1])
else
return 0
endA timeline example shows Client A's lock expiring, Client B acquiring it, then Client A waking up and deleting B's lock with a blind DEL, causing B to run unprotected. Redisson's RLock encapsulates UUID generation, Lua release, and watchdog renewal.
Pitfall 3: Lock Expires Before Business Finishes — GC Pauses / Network Hiccups
Even with proper TTL (e.g., 30s), a business operation taking longer (e.g., 35s due to RPC latency + 4s Full GC) will lose the lock while still executing. Another client acquires the lock and both write concurrently — data corruption. Martin Kleppmann's diagram illustrates this: Client 1 gets lease, enters stop-the-world GC, lease expires, Client 2 acquires lease and writes, Client 1 resumes and writes with stale lease. Checking the lock again before writing doesn't help because the check itself can be paused.
Lock TTL should be business P99 latency × 1.5–2, but this only reduces probability, cannot eliminate.
Pitfall 4: Watchdog Renewal Isn't a Silver Bullet — Main Thread Block Stops Renewal
Redisson's watchdog (default 30s TTL, renews every 10s) runs on Netty's HashedWheelTimer, which shares the EventLoop with business code. If the business thread blocks on synchronous I/O (old HttpClient, JDBC, SOAP) or experiences a long GC pause (>10s), the renewal task cannot run, the lock expires, and the scenario from Pitfall 3 repeats. The watchdog only proves the JVM is alive, not that business logic is progressing.
Correct practices:
Prohibit blocking calls in locked sections; use async (CompletableFuture, Reactor, Coroutines).
Set TTL = business P99 × 2.
Add idempotency at the resource layer (optimistic locking) as the second line of defense.
Pitfall 5: Master-Failover Loses Locks — Asynchronous Replication Flaw
In master-slave setups, a client acquires the lock on master, but before replication reaches the slave, the master fails. The slave promotes without the lock, allowing another client to acquire it. The first client, unaware, continues operating under the false belief it holds the lock. This is the fundamental flaw Redlock aims to solve by requiring a majority of independent nodes.
Mitigation: enable min-replicas-to-write 1 and min-replicas-max-lag 10 on the lock instance so master rejects writes if no healthy replica exists — a CAP tradeoff (consistency over availability).
Pitfall 6: Redlock's 5 Independent Nodes Assumption Rarely Holds
Redlock requires 5 truly independent Redis instances (separate data centers, power, NTP). Most teams deploy "5 Redis instances" in the same rack, same network, same NTP source — violating independence. Two hidden assumptions break safety:
Node clocks synchronized to milliseconds — any forward jump (manual NTP correction) causes premature expiry on that node.
Client never pauses longer than TTL — Full GC, SIGSTOP, scheduler preemption all violate this.
Antirez's 2016 rebuttal addressed network latency (algorithm measures total time T2-T1), clocks (only needs rough rate synchronization, promised monotonic clock — still not delivered as of 2026), and process pauses (acquisition pauses are safe; post-acquisition pauses affect all lock systems). The official Redis documentation still states: "Redis does not use a monotonic clock for TTL expiry; wall-clock jumps may cause a lock to be held by multiple processes."
90% of teams run Redlock on 5 nodes in one rack — the gap between algorithm assumptions and deployment reality is the biggest pitfall.
Pitfall 7: Clocks + GC Together Break Redlock — Fencing Tokens Are the Ultimate Guard
Combining master failover, clock jump, and GC pause yields Martin's "champion scenario": Client 1 gets majority, GC pauses 6s, Node 2's clock jumps forward 2s, Client 2 gets majority on other nodes, Client 1 resumes and both write — data corruption. The solution is fencing tokens : each lock acquisition returns a monotonically increasing token (e.g., etcd Revision, ZooKeeper zxid, DB version column). The resource rejects writes with tokens ≤ the highest seen token.
Client1: token=33 → GC 6s → writes with token=33 → rejected
Client2: token=34 → writes with token=34 → accepted
DB: UPDATE ... WHERE lock_token < 36Martin's second diagram shows the same timeline but with fencing token check at storage — the explosion (data corruption) disappears. Redis cannot natively provide fencing tokens because its keys are single-point; an INCR counter on one Redis doesn't synchronize across nodes. Therefore, Redis locks are inherently unsuitable for strict mutual exclusion .
Antirez counters: if storage can enforce token ordering, it's already a linearizable store and can generate its own IDs, making Redlock redundant; also, many locked resources (physical devices, external APIs) lack version columns. Both arguments are valid, but Martin's core point stands: for strict correctness, the lock service cannot be the only defense — the resource must enforce ordering.
Decision Matrix: Redis vs ZooKeeper vs etcd
Consistency model : Redis (AP), ZooKeeper (CP strong), etcd (CP strong + Raft)
Mutex primitive : Redis SET key uuid NX PX + Lua; ZooKeeper ephemeral sequential nodes; etcd Lease + PUT if Not Exists Fencing token : Redis ❌ (needs custom INCR, unsafe across nodes); ZooKeeper ✅ zxid built-in; etcd ✅ Revision built-in
Master failover loses lock : Redis ⚠️ Yes (async replication); ZooKeeper ✅ No (ZAB sync); etcd ✅ No (Raft)
Clock dependency : Redis ⚠️ TTL relies on local clock; ZooKeeper ✅ None; etcd ✅ None
Business exceeds TTL : Redis watchdog (on main thread → fails); ZooKeeper session heartbeat (separate thread → robust); etcd Lease KeepAlive (separate thread → robust)
Typical QPS : Redis 100k+ (single instance); ZooKeeper 10k; etcd 10k
Deployment barrier : Redis existing, near-zero cost; ZooKeeper needs dedicated ZK cluster; etcd existing or with K8s
When to use : Redis for efficiency, dedup, cache stampede; ZooKeeper for strong consistency, cross-DC, finance; etcd for K8s/cloud-native, cross-DC
Simplified: deduplication → Redis; cross-DC leader election → etcd or ZK; financial correctness → DB optimistic lock + Redis lock (lock optimizes, DB guarantees).
Author's Practical Guidelines
Efficiency scenarios (dedup, cron, cache stampede, short duplicate requests): single Redis + Redisson default RLock. TTL = P99 × 2, watchdog auto-renews.
Medium correctness (distributed scheduling, mild contention, retryable data): Redisson MultiLock (one client writes to multiple Redis), not Redlock.
Strong correctness (funds, inventory, order state machines): RLock + DB version column optimistic locking — dual safety.
Cross-DC / container orchestration : etcd concurrency package — native Revision + Lease + Raft.
Avoid entirely : Redlock's "5 independent nodes" deployment cost (5 independent racks + NTP + power) is prohibitive; most "5 Redis" deployments don't qualify.
Migration order: replace all SETNX+EXPIRE with SET key uuid NX PX; verify release logic checks UUID (use DELEX if supported, else Lua); then evaluate if 5-node Redlock is truly needed — answer is usually no.
Reliable distributed locking is always "lock service + resource-side version control" — the former saves cost, the latter guarantees data. Martin said this in 2016; in 2026 most production systems still repeat the same mistakes.
FAQ
Q1: Does SET NX PX really differ from SETNX+EXPIRE ? Can a single command crash between two commands? Yes. The two commands are separated by a network round-trip. If the client is OOM-killed, kill -9 ed, or network partitions for ~30s after SETNX succeeds, EXPIRE never reaches Redis. The key stays forever without TTL. SET NX PX is atomic — no exception.
Q2: My business takes 35s worst-case; should I set TTL to 35s? No. 35s is a worst-case outlier. Set TTL = P99 × 1.5–2. Better: make all remote calls async to push latency to milliseconds.
Q3: When does Redisson watchdog NOT work? Two boundaries: (1) explicit leaseTime passed (e.g., lock.lock(5, TimeUnit.SECONDS)) disables watchdog; (2) business thread blocks the shared EventLoop via synchronous I/O, preventing the timer task from running.
Q4: If I deploy Redlock on 5 Redis instances, are Pitfalls 6/7 solved? No. Pitfall 6 requires true fault-domain independence (separate racks, power, NTP) — same rack doesn't count. Pitfall 7 requires bounded process pauses and stable clocks — unguaranteed in any GC language/VM/container. Martin never retracted this; Redis docs still disclaim lack of monotonic clock.
Q5: How to choose etcd/ZK/Redis in production? If you have K8s, etcd is already there — use etcd clientv3 concurrency.NewMutex. Traditional finance/enterprise ops prefer ZK. If you have neither, don't force it — Redis lock + DB optimistic lock is what 80% of engineers actually use, and Martin doesn't oppose this.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
IT Services Circle
Delivering cutting-edge internet insights and practical learning resources. We're a passionate and principled IT media platform.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
