· via dev.to (home feed)
Claude Code agent's median run hit 79.8M tokens, 97% of them cache reads
A dev.to post logs 13 full runs of an autonomous Claude Code agent: a median 79.8M tokens each, about 97% cache reads and just 0.2% output — rare hard data for budgeting agent workloads.

A median run of 79.8 million tokens
A developer has published unusually concrete numbers on what an autonomous coding agent actually consumes: across thirteen complete runs of a Claude Code agent, the median run used 79.8 million tokens, and roughly 97% of that was cache reads — the same context, re-read on every response.
The post, published on dev.to by Rulestack, describes an agent that runs a small shop in public. It writes posts, publishes articles, replies to readers and checks its own health. The owner starts a full run with one short go-ahead word in Japanese, roughly meaning "do it," sometimes with an added note, and the agent works until every routine task on its list is either done or waiting on the owner.
How the ledger was built
Since 23 September 2026, a script has copied per-response usage out of Claude Code's session transcripts into a ledger, one row per prompt; when it was added, it also back-filled older sessions whose transcripts were still on disk. Each row records the model and effort level, the response count, and the input, cache-write, cache-read and output tokens, split between the main thread and subagents.
Thirteen rows in that ledger are full runs started by the go-ahead prompt, spanning 28 August to 1 October 2026. Totals varied widely: the smallest run used 12.3 million tokens, the largest 182.9 million.
Where the 80 million goes
The per-run breakdown, according to the post:
- Cache reads: median 97.3% of each run, ranging from 91.6% to 98.2%.
- Cache writes: median 2.5%.
- Output: median 167,540 tokens per run, about 0.2% of the total.
- Uncached input: a median of just 1,176 tokens per run.
- Subagents: median 29.9% of a run's tokens, ranging from 7.8% to 51.2%.
The totals are mostly a function of response count times context size. The median run had 248 responses in the main thread, and each of them re-read a median of roughly 218,000 tokens of cached context. The two biggest runs, at 181.6M and 182.9M, also logged the most main-thread responses (531 and 475) and the most subagent tokens (54.4M and 66.1M). In other words, 80 million tokens is not 80 million tokens of new work; it is roughly one long context read a few hundred times over, plus subagents that each start a context of their own.
What the data does not show
The author is explicit about the limits. They know what a run uses, but not what a run would lose if it used less: they have not tried splitting the routine into shorter runs with smaller contexts and comparing the results, and they have not separated the cost of reading project instructions from the cost of the work itself. The thirteen runs also changed models and effort levels over the five weeks, so they are not a controlled comparison — the largest run used the highest effort setting tried for full runs, but a medium-effort run reached 181.6M as well.
The post closes as a question to readers: how many tokens does one run of your agent use, and what even counts as a run — one prompt, one task, one day? If your cache-read share sits far below 97%, what breaks the cache in your setup? And has anyone deliberately shrunk a run, through shorter sessions, fewer subagents or a lighter CLAUDE.md, and did the results get worse?
Why it matters
Public, per-run token data for agents is rare; most teams only encounter it as a line on an invoice or as a rate-limit wall. The shape of this data is the useful part. Because output is a rounding error and almost everything is cache reads, raw token counts are a poor proxy for spend — cached input is billed at a steep discount relative to uncashed input on Anthropic's API — while the real cost drivers are context length, response count and subagent fan-out. Those are the levers to watch when budgeting long autonomous sessions. The post also demonstrates a cheap measurement method: Claude Code already writes per-response usage into its session transcripts, so a small script yields a full per-run ledger without instrumenting the agent itself, and the author is inviting others to publish comparable numbers.
- #claude-code
- #ai-agents
- #token-usage
- #prompt-caching
- #anthropic