deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Cloudflare's 2019 regex outage revisited: 27 minutes of 502s from one WAF rule

A dev.to retrospective revisits the July 2019 Cloudflare outage, when a regex in a simulate-mode WAF rule sent CPUs to 100% worldwide and caused 27 minutes of 502 errors.

Cloudflare's 2019 regex outage revisited: 27 minutes of 502s from one WAF rule

What happened on July 2, 2019

At 13:42 UTC on July 2, 2019, a newly merged firewall rule was deployed automatically to every Cloudflare machine across more than 180 cities. The rule contained a regular expression that, on ordinary traffic, sent CPU cores into a spiral of failed matching attempts. The first pager fired at 13:45, and sites behind Cloudflare began returning 502 errors worldwide. At the worst point traffic dropped by 82%, and the disruption lasted 27 minutes.

According to a dev.to retrospective, the recovery moved fast once the cause was found: the WAF was identified around 14:00, a global termination of the WAF was proposed at 14:02 and executed at 14:07, and traffic and CPU were normal again by 14:09. The WAF was switched back on at 14:52, without the offending rule. Cloudflare's CTO John Graham-Cumming published the full technical postmortem on July 12, ten days after a same-day summary from CEO Matthew Prince, who stressed that the incident was not an attack.

The regex at the center

The rule was meant to catch the inline JavaScript patterns used in cross-site scripting attacks. The full expression, published in the postmortem, is long, but Graham-Cumming singled out one fragment, .*(?:.*=.*), which reduces to .*.*=.*: match anything, then anything, then an equals sign, then anything. Two adjacent unbounded wildcards are the heart of the problem.

The rule was running in simulate mode, which records matches without blocking requests. That made no difference to the CPU bill: as the postmortem explained, rules still have to execute in simulate mode, so a rule that blocks nothing costs as much to run as one that blocks everything.

How catastrophic backtracking blows up

Cloudflare's WAF was written in Lua and matched with PCRE, a backtracking engine with no built-in protection against runaway expressions. A greedy .* first swallows the entire input, then gives back characters one at a time, and for every split of the first wildcard the engine also tries every split of the second. If an equals sign exists somewhere, the engine eventually finds it; if it does not, the engine must exhaust every combination before it can declare failure.

The postmortem's appendix quantifies this: 23 steps for x=x, 555 steps for x= followed by 20 x's, and 4,067 steps for 20 x's containing no equals sign at all. A trailing semicolon in the pattern pushed the last case to 5,353 steps. The worst case, a non-matching input, is also the most common case for a WAF scanning ordinary requests. When this behaviour is triggered deliberately it is known as ReDoS; here, nobody had to attack anything, because everyday traffic was enough.

The guards that were missing

The postmortem lists eleven contributing causes, and three stand out. A CPU limit that would have stopped runaway expressions had been removed by mistake weeks earlier, during a refactor that was itself intended to reduce WAF CPU usage. The test suite checked what rules blocked and allowed, but never measured how long a rule took to run. And WAF rules skipped the staged rollout process other Cloudflare software went through, deliberately, so that threats could be answered quickly: rules shipped through the Quicksilver key-value store, reaching every machine in about 2.3 seconds at p99. In the previous 60 days there had been 476 rule changes, roughly one every three hours, and this change was not an emergency. It still went global instantly.

A kill switch behind the outage

Even the recovery path was compromised. The internal control panel sat behind Cloudflare Access, which failed along with everything else, and some engineers' credentials had been disabled for infrequent panel use. Jira and the build system were unreachable too, and the team resorted to a rarely used bypass mechanism. Customers faced the same wall: the dashboard and API run through Cloudflare's own edge, so a control plane riding on the data plane fails with it.

What Cloudflare changed

The fixes included restoring the CPU protection, hand-reviewing all 3,868 managed WAF rules, adding performance profiling to the test suite, moving the WAF toward re2 or the Rust regex engine, both of which carry run-time guarantees, staged rollouts for non-emergency rules, and an emergency route for the dashboard and API that bypasses the edge.

Why it matters

The retrospective's central point is that nothing about this incident is historical. Anyone can write .*.*=.* today, and most code reviews and test suites will not catch it, because they test correctness rather than cost. The durable lessons are systemic: enforce hard limits on regex execution time, measure performance in CI, stage rule rollouts like any other code change, and never place the only kill switch behind the infrastructure it is meant to kill. The postmortem also set a standard for blameless analysis: its first cause is simply that an engineer wrote a regex prone to enormous backtracking, while the other ten describe the system that carried it to every server worldwide in seconds.

  • #cloudflare
  • #regex
  • #outage
  • #web-infrastructure
  • #waf

Related posts