· via Hacker News – Front Page (native)
celld's deterministic simulation testing catches alarm race in self-hosted Workers runtime
Self-hosted Cloudflare Workers runtime celld explains how deterministic simulation testing found and fixed an alarm race that could silently skip a confirmed scheduled task after a restart.

celld, a runtime for running Cloudflare Workers and Durable Objects applications on your own machines, has published an engineering write-up explaining how it uses deterministic simulation testing to hunt for concurrency bugs. The approach has already paid off: according to the post, the team's simulator surfaced a previously unknown race condition in alarm scheduling that could cause a confirmed alarm to be skipped after a node restart.
The write-up, which surfaced on Hacker News, describes celld as a distributed system with a deliberately small footprint — an S3-compatible object store is its only external service dependency. Distribution is also what makes bugs hard to chase: a failure might depend on a particular interleaving of delayed messages, failed writes and restarts that will not repeat on the next test run, which makes both diagnosis and verification of a fix difficult.
How the simulator works
According to celld, the simulator executes the runtime's production code inside a controlled environment. A cell — the unit that runs application code and owns a SQLite database — reacts to events such as incoming requests, completed storage operations and timer firings. The code that decides which event happens next is kept separate from the code that handles it, so the simulator can dictate ordering while the handling logic stays identical to what runs in production.
The simulator also controls when asynchronous tasks execute and how object storage responds, including delaying writes or making them fail, and it can advance simulated time without waiting for real time to pass. Tests define a space of allowed requests and faults; within those bounds, the simulator chooses randomly. A seed initializes the random number generator, so the same code, settings and seed always produce the same sequence of events and the same result. After every step, a checker validates invariants — conditions the system must uphold throughout a run, such as a confirmed alarm always keeping a way to wake its cell.
The alarm race it uncovered
Alarms in celld schedule a cell to run application code at a specified time, and idle cells can be unloaded from memory and reloaded when needed. Because opening every cell's SQLite database just to look for pending alarms would be expensive, celld maintains wake entries in object storage that record which cell to reactivate and when.
At the time, celld reused wake entries to cut down on storage writes: if an application deleted a 10:00 alarm and set a new one for 10:05, the lingering 10:00 entry would still wake the cell early enough. Deleting the old alarm also spawned a background cleanup task that removed the stale wake entry only after success had already been returned to the client, so the client would not wait on the extra storage operation.
The simulator found an ordering in which the client set the 10:05 alarm, received confirmation, and only then did the delayed cleanup run — deleting the single wake entry the new alarm depended on. The checker flagged the violated invariant: the alarm existed in SQLite, but nothing remained to reactivate the cell. If the node had restarted before 10:05, the alarm might never fire despite the earlier success response. Notably, celld says no hand-written test prescribed this sequence; the simulator discovered it through exploration.
Reproducing and fixing it
Because the same seed reproduces the failure exactly, engineers could inspect the precise moment the cleanup deleted the needed entry on every run without searching for the ordering again. The team has since redesigned wake entries so that each alarm setting writes its own entry in object storage. A delayed deletion for an older alarm can now only remove that older entry, leaving the new one intact, and stale entries are removed once celld confirms they are no longer needed. The simulator itself remains under development and is not part of celld's public repository.
Why it matters
Reproducibility is the scarcest resource in distributed systems debugging, and seeded simulation turns rare, nondeterministic failures into deterministic, replayable traces. The celld post is a compact demonstration of that payoff: a bug class — returning success before asynchronous cleanup finishes, combined with a reusable resource — that conventional example-based tests easily miss. As Workers and Durable Objects workloads move onto self-hosted runtimes outside Cloudflare's managed edge, correctness guarantees like alarms firing after restarts become the runtime operator's responsibility. For teams building similar infrastructure, the write-up doubles as a practical template: separate scheduling from handling, inject faults and time control at the boundaries, and let a checker assert invariants after every step.
- #cloudflare-workers
- #distributed-systems
- #testing
- #durable-objects
- #cloud