Skip to content
diff/reel
All reels
Databases

How Redis Caching Works

How Redis Caching Works — opening frame

sandboxed iframe · 82.5s loop · 39 KB

Made with Diffreel — draw your own →

A copy of the row, kept somewhere fast — 40× quicker to read, and wrong the moment the row changes.

By · Posted Aug 15, 2026 · 75 views

A cache is a copy of the answer, kept somewhere fast. The app asks the cache first; on a miss it reads the database and drops a copy on the way back, so the next read never travels that far. An in-memory read costs about 2 ms against roughly 80 ms for the query it replaces — illustrative figures, but the order of magnitude is the point.

Nothing updates that copy by itself. When the row changes, the cache keeps serving the old one: a stale cache. The standard fix is a TTL — every copy gets a timer and is thrown away at expiry, so the next read fetches the truth.

The timer then sets its own trap. When a popular key expires, every reader misses in the same instant and they all charge the database together — a cache stampede. Request coalescing closes it: only the first miss fetches, the rest wait for that one result and read the refilled copy.

The reel runs the whole ladder on one row: the copy lands, the copy lies, the timer fixes it, the timer causes a stampede, and one fetch rebuilds the shelf for everyone.

Source

How a Redis cache makes reads fast, and what the copy costs

4 min read

A Redis cache makes reads fast by keeping a copy of the answer in memory, and every problem it creates comes from that copy: the read is fast because it is a copy, and it is wrong for the same reason. The reel above climbs that ladder; this is the written version, with the formal names and the settings you have to choose.

The reel's stage: users, the app, the shelf, and the database.
The stage the reel argues on. Users talk only to the app; the app talks to the shelf above it and the database below it. There is no wire from a user to either store — only the app touches storage.

How a Redis read-path cache fills

The app answers a read by asking Redis first. If the key is there, that is the answer and nothing else runs. If it is not — a miss — the app reads the database and writes the copy into Redis on the way back.

That last clause is the whole mechanism: the read path is also the write path for the cache, and a miss is what populates it. The shape is called cache-aside when the application owns the fetch and read-through when the cache library does; the sequence is the same either way.

One read, two places
Database read80 ms
Cache read2 ms

Illustrative and order-of-magnitude correct, not telemetry: Redis serves reads in well under a millisecond to low single digits, a relational query commonly costs tens. The reel's 40x is 80/2, derived on stage from these two.

Here is the whole read path with every fix already in it. The sections below take it apart in the order the reel does.

ts
const TTL_SECONDS = 300; // 5:00

async function getUser(id: string) {
  const key = `user:${id}`;

  const cached = await redis.get(key);
  if (cached) return JSON.parse(cached); // hit: ~2 ms, the database never hears about it

  // singleFlight stands in for whatever coalescing primitive the stack provides —
  // it is not a Redis client call
  return singleFlight(key, async () => {
    // only the first miss for this key runs this block; the rest await its result
    const row = await db.query("select * from users where id = $1", [id]); // ~80 ms
    await redis.set(key, JSON.stringify(row), { EX: TTL_SECONDS });
    return row;
  });
}

Why a cached copy goes stale

Nothing in that read path updates the copy. Change the row — a user renames themselves from Alan to Alana — and there are two answers to one question. The database holds the new one, Redis holds the old one, and the app reads Redis.

This is a stale cache, and it is the defining failure of a naive read-path cache rather than a bug in any implementation. The dangerous property is not that the copy is wrong; it is that the window is unbounded. Nothing removes it, nothing reports it, and the first signal is usually a user who can see their own old name.

TTL strategy: how long the copy may lie

The standard fix is a time-to-live — TTL — on every entry. Five minutes, say. The timer runs out, the entry is evicted, the next read misses, and the miss refetches.

Be precise about what that buys. A TTL does not make the cache correct; it converts unbounded staleness into a bounded number you chose. A TTL strategy is the reasoning behind that number, not the number: how stale this particular data may be, how expensive the refill is, and whether a whole class of keys will expire together. Shorter means fresher and more database reads; longer means cheaper and a longer window in which the app is confidently wrong. There is no setting that is both.

The cache stampede a TTL sets off

Take a popular key that a thousand concurrent readers want. While the copy is on the shelf the database sees no reads for it. Then the timer expires, all thousand miss at once, every miss means "go and read the database", and a thousand queries arrive together for a row that was costing nothing.

Database reads for one hot key, across its expiry
Naive TTLWith coalescing
1,0005000
Naive TTLWith coalescing1,0001
t-2st-1sexpiryt+1s

The reel's scenario: 1,000 readers at the moment one hot key expires.

This is a cache stampede, also called a thundering herd. What makes it bite is the flat line either side of the spike: incoming traffic never changed. The cache took down the database it was there to protect.

Request coalescing: only the first miss fetches

Stop treating a thousand simultaneous misses as a thousand requests. Only the first miss for a key reaches the database; the other 999 wait on that in-flight fetch and read the copy it leaves behind.

The coalesced refill
  1. App

    1,000 reads arrive for user:42, just after the copy expired

  2. Cache

    miss — the key isn't there

  3. App

    the first miss goes through; the other 999 wait on it

  4. Database

    one query, ~80 ms, one row

  5. Cache

    the copy lands again, with a fresh 5:00 timer

  6. App

    the 999 wake and read the copy — 2 ms each

Request coalescing: a thousand misses collapse into a single database read, and the refill answers all of them.

This is request coalescing, or single-flight. The 999 waiters still pay the latency of the one fetch, roughly 80 ms — but the database's exposure at the expiry moment drops from a thousand reads to one.

Every fix on this ladder except the last creates the failure the next one answers. That is not a defect in caching; it is what caching is.

Redis cache eviction policy vs TTL

Expiry and eviction are different mechanisms, and confusing them is the usual production surprise. A TTL removes one key when its own timer runs out. A Redis cache eviction policy — maxmemory-policy — decides which keys go when the instance hits its memory ceiling, whatever their timers say.

The default is noeviction, which frees nothing and starts failing writes. allkeys-lru evicts the least recently used key from the whole space. The volatile-* policies only consider keys that have a TTL, so an instance where nobody sets TTLs falls back to erroring on write.

Redis vs Memcached

Memcached is a multithreaded LRU cache of opaque blobs, and it is very good at exactly that. Redis is single-threaded per command, with data structures, replication, persistence and the eviction policies above — more surface, and more to reason about. "Redis because we already run Redis" is an honest answer; "Redis is faster" is not.

Redis caching in a system design interview

Three moves separate a strong answer from a recited one: name the stale window before you are asked about it, give a TTL strategy rather than a TTL, and say out loud that eviction is not expiry. Reach for coalescing on "hot key", and TTL jitter on "everything expired at once" — siblings, not the same fix.

What this leaves out

The ladder stops at coalescing because that is one mechanism per failure. These neighbours are real and not covered here: TTL jitter, for many keys expiring together rather than many readers missing one key; stale-while-revalidate, serving the expired copy while the refresh runs; cache warming; write paths (cache-aside versus write-through, a different reel); and hot-key replication.

What the reel claims

Every number and assertion on the canvas, and where it comes from.

On-stage claimVerification
DB read ~80 ms, cache read ~2 ms, 40×Illustrative ballpark; 40 = 80/2, derived on stage; Redis sub-millisecond to low-millisecond reads are documented
The copy is written on the way back from a missThe cache-aside / read-through read path, in all sources
The copy goes stale when the database row changesDefinitional for a read-path cache with no invalidation
A TTL evicts the copy; the next read refetchesStandard TTL semantics
A hot key expiring sends every reader to the database at onceCache stampede / thundering herd
Only the first miss fetches; the rest wait and read the copyRequest coalescing / single-flight

Sources

Related reels