How a Redis cache makes reads fast, and what the copy costs
4 min read
A Redis cache makes reads fast by keeping a copy of the answer in memory, and every problem it creates comes from that copy: the read is fast because it is a copy, and it is wrong for the same reason. The reel above climbs that ladder; this is the written version, with the formal names and the settings you have to choose.

How a Redis read-path cache fills
The app answers a read by asking Redis first. If the key is there, that is the answer and nothing else runs. If it is not — a miss — the app reads the database and writes the copy into Redis on the way back.
That last clause is the whole mechanism: the read path is also the write path for the cache, and a miss is what populates it. The shape is called cache-aside when the application owns the fetch and read-through when the cache library does; the sequence is the same either way.
Illustrative and order-of-magnitude correct, not telemetry: Redis serves reads in well under a millisecond to low single digits, a relational query commonly costs tens. The reel's 40x is 80/2, derived on stage from these two.
Here is the whole read path with every fix already in it. The sections below take it apart in the order the reel does.
const TTL_SECONDS = 300; // 5:00
async function getUser(id: string) {
const key = `user:${id}`;
const cached = await redis.get(key);
if (cached) return JSON.parse(cached); // hit: ~2 ms, the database never hears about it
// singleFlight stands in for whatever coalescing primitive the stack provides —
// it is not a Redis client call
return singleFlight(key, async () => {
// only the first miss for this key runs this block; the rest await its result
const row = await db.query("select * from users where id = $1", [id]); // ~80 ms
await redis.set(key, JSON.stringify(row), { EX: TTL_SECONDS });
return row;
});
}Why a cached copy goes stale
Nothing in that read path updates the copy. Change the row — a user renames themselves from Alan to Alana — and there are two answers to one question. The database holds the new one, Redis holds the old one, and the app reads Redis.
This is a stale cache, and it is the defining failure of a naive read-path cache rather than a bug in any implementation. The dangerous property is not that the copy is wrong; it is that the window is unbounded. Nothing removes it, nothing reports it, and the first signal is usually a user who can see their own old name.
TTL strategy: how long the copy may lie
The standard fix is a time-to-live — TTL — on every entry. Five minutes, say. The timer runs out, the entry is evicted, the next read misses, and the miss refetches.
Be precise about what that buys. A TTL does not make the cache correct; it converts unbounded staleness into a bounded number you chose. A TTL strategy is the reasoning behind that number, not the number: how stale this particular data may be, how expensive the refill is, and whether a whole class of keys will expire together. Shorter means fresher and more database reads; longer means cheaper and a longer window in which the app is confidently wrong. There is no setting that is both.
The cache stampede a TTL sets off
Take a popular key that a thousand concurrent readers want. While the copy is on the shelf the database sees no reads for it. Then the timer expires, all thousand miss at once, every miss means "go and read the database", and a thousand queries arrive together for a row that was costing nothing.
The reel's scenario: 1,000 readers at the moment one hot key expires.
This is a cache stampede, also called a thundering herd. What makes it bite is the flat line either side of the spike: incoming traffic never changed. The cache took down the database it was there to protect.
Request coalescing: only the first miss fetches
Stop treating a thousand simultaneous misses as a thousand requests. Only the first miss for a key reaches the database; the other 999 wait on that in-flight fetch and read the copy it leaves behind.
App
1,000 reads arrive for user:42, just after the copy expired
Cache
miss — the key isn't there
App
the first miss goes through; the other 999 wait on it
Database
one query, ~80 ms, one row
Cache
the copy lands again, with a fresh 5:00 timer
App
the 999 wake and read the copy — 2 ms each
Request coalescing: a thousand misses collapse into a single database read, and the refill answers all of them.
This is request coalescing, or single-flight. The 999 waiters still pay the latency of the one fetch, roughly 80 ms — but the database's exposure at the expiry moment drops from a thousand reads to one.
Every fix on this ladder except the last creates the failure the next one answers. That is not a defect in caching; it is what caching is.
Redis cache eviction policy vs TTL
Expiry and eviction are different mechanisms, and confusing them is the usual production surprise. A TTL removes one key when its own timer runs out. A Redis cache eviction policy — maxmemory-policy — decides which keys go when the instance hits its memory ceiling, whatever their timers say.
The default is noeviction, which frees nothing and starts failing writes. allkeys-lru evicts the least recently used key from the whole space. The volatile-* policies only consider keys that have a TTL, so an instance where nobody sets TTLs falls back to erroring on write.
Redis vs Memcached
Memcached is a multithreaded LRU cache of opaque blobs, and it is very good at exactly that. Redis is single-threaded per command, with data structures, replication, persistence and the eviction policies above — more surface, and more to reason about. "Redis because we already run Redis" is an honest answer; "Redis is faster" is not.
Redis caching in a system design interview
Three moves separate a strong answer from a recited one: name the stale window before you are asked about it, give a TTL strategy rather than a TTL, and say out loud that eviction is not expiry. Reach for coalescing on "hot key", and TTL jitter on "everything expired at once" — siblings, not the same fix.
What this leaves out
The ladder stops at coalescing because that is one mechanism per failure. These neighbours are real and not covered here: TTL jitter, for many keys expiring together rather than many readers missing one key; stale-while-revalidate, serving the expired copy while the refresh runs; cache warming; write paths (cache-aside versus write-through, a different reel); and hot-key replication.
What the reel claims
Every number and assertion on the canvas, and where it comes from.
| On-stage claim | Verification |
|---|---|
| DB read ~80 ms, cache read ~2 ms, 40× | Illustrative ballpark; 40 = 80/2, derived on stage; Redis sub-millisecond to low-millisecond reads are documented |
| The copy is written on the way back from a miss | The cache-aside / read-through read path, in all sources |
| The copy goes stale when the database row changes | Definitional for a read-path cache with no invalidation |
| A TTL evicts the copy; the next read refetches | Standard TTL semantics |
| A hot key expiring sends every reader to the database at once | Cache stampede / thundering herd |
| Only the first miss fetches; the rest wait and read the copy | Request coalescing / single-flight |
Sources
- Redis — Key eviction (
maxmemory-policy, the eight policies,noevictionas the default) - Redis — How to tame the thundering herd problem
- OneUptime — How to handle cache stampede in Redis
- OneUptime — Cache stampede prevention
- VergeCloud — What is cache stampede?





