· via Cloudflare blog
Cloudflare adds on-demand CPU and memory flamegraph profiling for Workers
Cloudflare now lets developers capture CPU and memory profiles of production Workers and Durable Objects and inspect them as interactive flamegraphs, from the dashboard or the CLI.

On-demand profiling goes live
Cloudflare has added on-demand CPU and memory profiling for Workers and Durable Objects, with results rendered as interactive flamegraphs. According to the Cloudflare blog, developers can request a profile of a live Worker from the Observability page in the dashboard, view it as a flamegraph, and download the profile file for deeper offline analysis.
Two ways to capture a profile
There are two entry points. In the dashboard, you navigate to Build, then Compute, then Workers & Pages, select your Worker and open its Observability tab, where a dropdown labelled Flamegraph offers both CPU and memory profiles. Alternatively, with the cf package installed, a single command captures a five-second CPU profile to a file:
cf workers versions profile latest
--worker-id "$WORKER_ID_OR_NAME"
--duration-ms 5000
--profile-type cpu > worker-cpu.pprof
The duration setting controls how long the profiler runs, and you can choose which version of the Worker to target. Cloudflare cautions that the Worker needs a healthy amount of traffic for a capture to succeed, so quiet versions are poor candidates.
Reading the results
Each rectangle in the flamegraph represents a function call, with its width showing how much CPU time or memory that function consumed. Clicking a function focuses the view, hovering reveals detail, and a table view ranks the functions most frequently seen in the sample. Cloudflare's advice is to take a few profiles and hunt for the widest boxes. One gotcha: TypeScript projects should enable source maps, otherwise the flamegraph will show obfuscated function names.
Two internal case studies
Cloudflare says several internal teams have already used the feature to optimise resource usage and fix out-of-memory errors, and the blog details two examples.
The first concerns the Worker implementing the R2 binding, profiled for 50 seconds. The wide boxes included expected costs such as decryptBlock and fillResponse, but sorting the table view by samples flagged genericR2JsonReplacer at over 5% of CPU time. The function was invoked by JSON.stringify for every node in a JSON tree while also walking the tree itself, so a value nested five levels deep was processed five times. Removing that duplicate work made the function 2.7x faster. The same profile exposed a redundant call to registry.metrics(), where a single invocation cost 1% of CPU time; caching the result in a variable eliminated the waste.
The second case was an internal Worker whose P999 memory sat around 133 MB against the 128 MB Worker limit, triggering frequent "Exceeded Memory" evictions. Errors and metrics confirmed memory was the problem but could not say which code caused it. A heap profile opened in pprof showed that the Worker's Prometheus instrumentation accounted for roughly 66.7% of allocations — code the team assumed was disabled, but which was only partly switched off and still paying the full memory cost. Deleting it entirely produced:
| Percentile | Before | After |
|---|---|---|
| P50 | 70 MB | 54 MB |
| P90 | 94 MB | 79 MB |
| P99 | 113 MB | 97 MB |
| P999 | 133 MB | 118 MB |
That left roughly 10 MB of headroom under the limit at P999, and the "Exceeded Memory" errors dropped accordingly.
Why profiling at the edge is hard
Workers have supported local profiling through Chrome DevTools for some time, but a local session never sees the same kinds or volume of requests as production. Production profiling at the edge adds complications: Workers are replicated across data centres and physical servers to keep requests near clients, and Durable Objects can be placed dynamically and can move. Before starting a session, the system must work out which version to profile, which data centre recently ran it, whether the isolate is still loaded there, whether it belongs exclusively to the requesting account, and, for a Durable Object, where the live primary actor resides. The client making the request, usually the dashboard, supplies some of these answers.
Why it matters
Serverless debugging has long leaned on logs and aggregate metrics, which can tell you a Worker is exceeding its memory limit or burning CPU, but not which function is responsible. Production flamegraphs close that gap, and because Workers enforce hard memory and CPU ceilings, the payoff is concrete: fewer evictions, better tail latency and less waste. The internal examples make the point — neither the recursive JSON replacer nor the half-disabled Prometheus path would have been obvious from dashboards alone. For teams running Workers at scale, this turns resource debugging from guesswork into a targeted lookup.
- #cloudflare-workers
- #serverless
- #profiling
- #debugging
- #observability