· via dev.to (home feed)
Postmortem: deploy-time cache warming made 180,000 keys expire at once
A catalogue service's cache-warming fix gave every entry the same one-hour TTL, so an hour after each deploy the whole cache expired together and saturated the database. Four changes, including TTL jitter, fixed it.

What the team saw
According to a postmortem by Sergey Shinder on dev.to, a product catalogue service began suffering a roughly six-minute degradation starting in May, arriving about an hour after every deployment. Database CPU reached its ceiling, page load times climbed from under a second to eight or nine seconds, and then the system healed on its own. On some days a smaller second wave followed an hour after the first. The engineers reportedly spent two days hunting for an hourly cron job that did not exist.
The improvement that caused it
The trigger was a change the team had been happy about. Newly deployed pods used to start with an empty cache and run slowly for the first few minutes while real traffic filled it in. To eliminate that cold start, they added a warm-up stage: before a pod took any traffic, it preloaded roughly 180,000 product records into the shared cache, each with a one-hour time to live. Cold starts disappeared, and a hidden problem was manufactured at the same time.
Because the warm-up ran for about two minutes, every entry was created inside the same two-minute window, and an hour later the entire catalogue expired inside the same window. Every request missed, every miss went to the database, and the database was provisioned for the small percentage of misses a steady-state cache produces. The second wave had the same origin: entries refilled during the surge were refilled together, so they expired together again. The lazy cache the team had replaced never behaved this way, because ordinary traffic had scattered expiry times across the hour without anyone deciding to.
Four mitigations
Shinder describes four changes that resolved the incidents:
- Every TTL now carries twenty percent random jitter, so entries loaded together expire apart.
- Misses on the same key are coalesced: a thousand simultaneous requests for one product result in one database query while the rest wait.
- Entries are served stale for up to five minutes past expiry while a single background refresh replaces them, making expiry invisible to users.
- Cache fills run through their own small pool of database connections, so the worst a stampede can do is slow the refill rather than take connections away from everything else.
Why it matters
The core lesson is about synchrony that nobody designed. Systems often spread load over time by accident, and that accident can quietly be load-bearing. A cleanup or optimisation that looks obviously safe, such as pre-populating a cache at deploy, can strip away randomness that was functioning as protection, and the failure only shows up an hour later, far from the change that caused it. The debugging shape is also instructive: recurring, self-healing incidents that look cron-shaped may in fact be self-inflicted synchrony from your own deploy-time behaviour.
For anyone building deploy pipelines that pre-populate caches, or batch-insert anything time-keyed, the practical checklist is short. Add jitter to every TTL by default. Coalesce misses on hot keys. Serve stale data behind a background refresh so expiry is not user-visible. Isolate cache-fill traffic on its own connections so a refill storm cannot starve the rest of the system. And measure the distribution of expiry times, not just hit rates, because a cache with a 99 percent hit rate can still take a database down if all the misses land in the same minute.
- #caching
- #databases
- #postmortem
- #deployment
- #reliability