Skip to content
diff/reel
All reels
Distributed Systems

Retry Storm vs Jitter

Retry Storm vs Jitter — opening frame

sandboxed iframe · 45s loop · 26 KB

Made with Diffreel — draw your own →

Same clients, same backoff, same traffic — one number changes, and the pile-on disappears.

By · Posted Aug 15, 2026 · 3 views

A service fails. Every client notices at the same moment, waits the same computed delay, and retries — together. The retry wave is synchronized, which means the recovering service is hit by its whole client base at once, fails again, and the cycle repeats.

Jittered backoff changes one thing: the delay is drawn at random below the exponential ceiling instead of being computed exactly. Same clients, same ceiling, same total number of retries — spread across the window instead of stacked on one instant.

The animation holds everything constant except comet timing. That is the entire lesson: the fix is not fewer retries, it is unsynchronized ones.

Source

Retry Storm vs Jitter: why exponential backoff alone won't save you

4 min read

A retry storm and jittered backoff are the same clients making the same number of retries against the same service — the only thing that changes is timing, and timing is the whole fight. In a retry storm every client that failed at the same instant computes the same backoff delay and wakes at the same instant, so all of them slam the recovering service together. Jittered backoff draws each client's wait at random from the same interval, so the identical set of retries arrives spread across the window instead of stacked on one edge. Put plainly, the difference between a synchronized retry storm and jittered backoff is not how often clients retry — it is whether they retry in phase.

Three client nodes on the left connected to one server node on the right.
The reel's stage: three identical clients on the left, one server on the right. In both scenes the same three retries fly; only their timing changes — a wave in one scene, a trickle in the other.

The wave: a synchronized retry storm

When a shared service fails, it fails for everyone at once. Every client learns of the outage in the same window, so unless their retries are deliberately decorrelated, they come back in the same window too. Plain exponential backoff does not decorrelate them. Each client computes its next wait as min(cap, base × 2^attempt) — a deterministic formula. Two clients that failed together and are on the same attempt number compute the same number, wake at the same instant, and fire together.

Backoff lowered the height of each wave; it did nothing to the synchronization. The wave hits a service that has just come back with cold caches and empty connection pools, re-kills it, and the cycle repeats one backoff step wider. The failure is not that clients retry too often — it is that they retry in phase. This is the classic thundering herd problem: a spike of identical requests landing at exactly the same moment, which is precisely how a retry storm turns a brief blip into a sustained outage.

The trickle: jittered backoff

Jitter here means deliberately injected randomness in the delay — not the networking sense of the word, where jitter is unwanted variance in packet latency and a bad thing. Here it is the fix. The exponential ceiling grows exactly as before; the wait is simply drawn from underneath it: random(0, min(cap, base × 2^attempt)). Two clients on the same attempt now compute different numbers, so the herd is broken on the very first retry rather than eventually. The same total number of retries arrives as a trickle spread across the window, and the service serves them.

This pattern — exponential backoff with jitter — is what mature retry libraries default to, and the AWS SDKs adopt it out of the box. AWS's own simulation with 100 contending clients found that full jitter reduced the call count by more than half and finished sooner, while un-jittered backoff was the clear loser on both work and time. That is the counter-intuitive part: adding randomness made the system finish faster, not slower.

Peak concurrent retries hitting the server (reel's 3-client scene)
Synchronized retry storm3 requests at once
Jittered backoff1 requests at once

Same three clients, same total retries — only the timing differs. Peak-3 vs peak-1 is the reel's measured deliverable.

When to add jitter to your backoff

Add jitter whenever more than one client retries the same dependency — which, at any real scale, is always. The moment you have a fleet of workers, a pool of app servers, or a stampede of mobile clients all talking to one backend, a bare sleep(2^attempt) is a loaded gun: the first shared failure synchronizes them, and every wave after it stays in phase.

Full jitter — random(0, ceiling) — is the simplest form and the safe default. If you want a floor under the delay so no client retries instantly, equal jitter (half + random(half)) or decorrelated jitter keep some spacing while still breaking the phase. The one thing you should never ship is exponential backoff with no randomness at all. Jitter costs a single call to your random number generator; the storm it prevents costs an outage.

Retry Storm vs Jitter in a system design interview

Propose retries in a design interview and the next question is almost always: "what stops those retries from taking the service down?" The trap answer is "exponential backoff" — a trap because backoff is necessary but not sufficient. It reduces how hard each client hits; it does nothing to stop N clients from hitting together.

The crisp answer names two mechanisms. First, jitter: randomize the delay so failed clients decorrelate instead of synchronizing into a thundering herd. Second, a cap on total retry volume — a retry budget or token bucket, and ideally a circuit breaker that stops retries entirely once a dependency is clearly down. If you can say "backoff sets the rate, jitter breaks the synchronization, and a retry budget bounds the total," you have answered the question the retry storm is really asking — and shown you know that the randomness, not the backoff, is the load-bearing half.

Sources

Related reels