deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Firebase postmortem: two-line flag cleanup crashed iOS apps while status dashboards stayed green

A Firebase postmortem explains how deleting one stale remote-config flag produced a malformed payload that crashed iOS apps on launch for over two hours while status dashboards stayed green.

Firebase postmortem: two-line flag cleanup crashed iOS apps while status dashboards stayed green

What happened

A routine housekeeping change in Firebase's remote-config system took down iOS applications around the world on September 28, and Firebase explained it publicly four days later, on October 2. According to the company's postmortem, at 17:38 Pacific Time a cleanup removed a stale legacy remote-config flag. The iOS SDK still held a reference to the deleted flag, and when clients went to fetch it, the lookup produced a fatal error. Within roughly three minutes, the resulting malformed configuration payload had reached Firebase's global customer base, and affected apps began crashing on launch.

Mobile teams use remote-config flags as kill switches: a way to disable a faulty feature without shipping a new app binary. As a dev.to analysis of the report points out, that makes the failure mode particularly awkward. The outage came not from the features the flags were guarding, but from housekeeping on the flags themselves — a change, the report reportedly notes, of just two lines.

The timeline, to the minute

The postmortem, as summarized by dev.to, lays out the incident clock:

  • 17:38 — the stale flag is removed
  • 17:41 — the malformed payload rolls out globally; clients crash on launch
  • 17:59 — crash alerts spike and the first external bug reports land on GitHub; the on-call engineer is paged
  • 19:16 — the culprit change is identified and a rollback begins
  • 19:52 — the rollback is fully deployed

That works out to detection in under twenty minutes, and two hours and eleven minutes between the bad payload starting to serve and full mitigation. Elevated error rates persisted well beyond that window, which the report attributes to lag in client-side error reporting.

Green dashboards, crashing phones

The detail that made the incident notable beyond Firebase's customer base is that the status dashboards never turned red. The postmortem explains why: the dashboards track server-side metrics, and the servers stayed healthy. The failure lived entirely on the client, where the dashboards were not looking.

According to the dev.to write-up, updating the status page required manual intervention that took hours, which meant the most current public record of the outage during the event was a GitHub issue thread. The dev.to commentary argues that this communication gap, more than the config bug itself, is the finding that deserves top billing in the retrospective.

What the report leaves out

The dev.to analysis also flags what the postmortem does not say. The blast radius is described only as a large number of affected iOS applications — no count of apps, devices or SDK versions. Third-party coverage put the figure in the thousands, but that number came from reporters tallying crash reports rather than from Google, which could read the exact total from its own logs. The report likewise states that existing tests passed on the change without naming the suite or describing what it covered, leaving other teams unable to judge whether their own tests would have caught the same fault.

The fix

Firebase has committed to an SDK patch that makes missing or corrupted flags fall back to cached defaults instead of crashing. As of the dev.to piece it was promised within a week, which means every app on the affected SDK line carried the same latent fault until it shipped. The company's own stated lesson is that pushing configuration changes quickly to a broad slice of its global customer base introduces unnecessary risk — the implication being that flags need phased rollouts and canaries at least as much as binaries do, since there is no build step to catch a bad one.

Why it matters

Two lessons generalize well past Firebase. First, safety mechanisms have lifecycles too, and the cleanup, migration or refactor of a control almost never inherits the control's own tests. Kill switches, circuit breakers and backups are all exposed to exactly this pattern: the configuration that breaks your system can be the configuration that was supposed to save you. Second, a status page that watches only servers will stay green through any client-side outage. If you ship client software and collect no client-side signal — crash reporting, a canary app, a third-party tracker — your next incident could look identical: healthy dashboards, crashing apps, and the real story surfacing on GitHub.

  • #firebase
  • #ios
  • #remote-config
  • #postmortem
  • #cloud
  • #incident-management

Related posts