deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Coordinated omission: why closed-loop load tests understate your p99 latency

Closed-loop load generators skip the requests that would reveal worst-case latency during stalls, making reported p99 numbers a floor rather than a true tail measurement.

Coordinated omission: why closed-loop load tests understate your p99 latency

A recent article on dev.to explains why the percentile numbers produced by most load-testing tools systematically understate tail latency. The phenomenon is called coordinated omission, and according to the piece it can leave a service that reported a 12 ms p99 in load testing exhibiting multi-second stalls for real users — while dashboards built on the same percentile math continue to look healthy.

How the bias works

Most load generators run as closed loops: send a request, wait for the response, record the latency, then send the next one. The article cites ab, naive wrk scripts, and many homegrown harnesses as tools that work this way by default. It feels like a faithful measurement, because every recorded sample is the real duration of a real request.

The trouble shows up when the system stalls — a garbage-collection pause, a lock-contention spike, or a downstream timeout. In the article's example, a closed-loop generator targeting 100 requests per second hits a one-second stall. During that second it has exactly one request in flight; that request returns after 1000 ms and becomes a single bad sample. Real traffic at 100 req/s, however, would have delivered roughly 100 requests in the same interval, all queuing behind the stalled one and each suffering close to a full second of delay. The generator never issued those requests because it was blocked waiting for a reply, so they never enter the percentile calculation at all.

That is the coordinated part of the omission: the missing samples vanish precisely when latency is at its worst — exactly the region a p99 exists to describe.

The scale of the distortion

According to the dev.to article, the term originates with Gil Tene of Azul Systems, creator of HdrHistogram, who first applied it to JVM pause-measurement tools that misrepresented the impact of garbage collection. The concept generalises to any closed-loop benchmark of any system: HTTP APIs, databases, queues, and inference endpoints.

The magnitude is not a rounding error. The article reports that recomputing percentiles on an arrival-rate basis — effectively accounting for the omitted requests — routinely turns a reported single-digit-millisecond p99 into several hundred milliseconds, or up to a full second. The true tail can be roughly ten times worse than the figure on the report, concentrated in the exact part of the distribution that tail percentiles are meant to expose.

What to do about it

The piece offers three fixes.

First, use an open-loop or corrected load generator. Tools such as wrk2, k6 with arrival-rate executors, Gatling, and Locust's constant-arrival-rate shape issue requests on a fixed schedule regardless of whether earlier requests have returned, which mirrors how production clients behave — real users do not queue behind your slowest request.

Second, if switching tools is impractical, log when each request should have been sent versus when it actually was, and compute corrected percentiles from the difference. HdrHistogram includes built-in methods for this kind of coordinated-omission correction.

Third, reinterpret every closed-loop percentile as a lower bound. A test reporting a 12 ms p99 supports only the claim that p99 is at least 12 ms under ideal, non-stalled conditions — not that p99 is 12 ms. That reframing alone changes how much confidence a passing load test deserves before a launch.

Why it matters

The article argues the problem is becoming more acute, not less, for LLM-backed services. Inference endpoints have highly variable per-request latency — cache hits versus cold starts, short versus long generations — so stalls are frequent rather than exceptional. A closed-loop load test on such a system actively conceals worst-case behaviour at the moment you most need to see it. For any team using load-test percentiles as a release gate, the practical takeaway is to check whether the tool measures latency at a fixed arrival rate, and to treat closed-loop p99 figures as floors rather than facts.

  • #load-testing
  • #latency
  • #benchmarking
  • #performance
  • #cloud

Related posts