deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Claude Code cache hit rates can hide 20M wasted tokens a day from cron jobs

A dev.to analysis shows how an account-level cache hit rate averaged away cron jobs rewriting the same 30K-token prefix every run, and details per-caller fixes that cut one bill by 13.6%.

Claude Code cache hit rates can hide 20M wasted tokens a day from cron jobs

A dev.to post argues that the cache hit rate on a Claude Code dashboard can look healthy while hiding serious waste, because a single account-level number averages over callers that share nothing. The opening example: a deployment reporting a 77% cache hit rate overall, while cron tasks on the same account ran at roughly 0%, because every scheduled run opens a fresh session and rewrites the same static prefix of about 30,000 tokens. At 672 runs per day, that is roughly 20 million input tokens daily that never touch a cache — numbers an operator posted to an open issue and the author quotes.

How prompt caching is priced

Anthropic prices caching with multipliers rather than a flat discount, and the post lays out the arithmetic. A write with a 5-minute TTL costs 1.25× the base input price, a 1-hour write costs 2×, and reads cost 0.1× on most models (the example table lists 0.05× on Opus 5.5 and 0.025× on Fable 5.1). Three consequences follow: a single read makes a 5-minute write worthwhile (1.35× total versus 2× for two uncached passes), the TTL should match the gap between calls, and a breakpoint nobody reuses costs 25% to 100% extra instead of saving anything. Every hit renews the TTL at read price, so a busy cache persists indefinitely, and the keepalive crossover lands at 62.5 minutes on every model because it is a ratio, not a model property.

One account, several cache behaviours

According to Anthropic documentation cited in the post, Claude Code on subscription billing within plan usage gives the main conversation a 1-hour TTL and everything else 5 minutes; on an API key, via a cloud provider, or past plan limits, everything gets 5 minutes. Subagents, workflows, forks, compaction calls and session titles all fall into that second bucket, and overrides exist from version 2.1.242 onward.

The structural detail is that a subagent's first request cannot read the parent's cache, because a different prompt and tool set makes the prefixes diverge at the first token. A fork, by contrast, inherits the parent's prefix exactly and hits on its first request.

The post also flags a discrepancy: the docs say subagents get 5 minutes, one user's transcripts showed every subagent write billed in the 1-hour bucket, and a maintainer said the effective TTL is decided further down the pipeline while the docs are corrected. The author's advice is to trust the raw response fields ephemeral_5m_input_tokens and ephemeral_1h_input_tokens rather than any summary table.

Why blanket 1-hour TTL backfires

The post's measured failure mode involves a parent that dispatches a subagent and then waits. No requests go out, so nothing refreshes the parent's cache. The median child runtime was about 9 minutes, just past expiry, and 96% of those waits ended with the cache genuinely expired: when the parent resumed, at least half of its cached prefix had to be rewritten at full price.

The counterintuitive result is that a blanket 1-hour TTL made the bill 8.6% worse. Reuse almost never needs the long window: 98% of cache hits landed within about 34 seconds of the write, with a median gap of 7 seconds. Targeted changes did pay off — a 1-hour write on the dispatch turn cut the bill 6.0%, a persistent per-type static prefix 1.0%, and moving dynamic content after the stable prefix 7.6%, for a combined 13.6%.

Where caching fails silently

Several failure modes produce no errors at all. Token floors: current-generation models cache at a 512-token minimum, while Opus 4.6/4.5 and Haiku 4.5 need 4,096, and below the floor the cache-creation field simply reads zero. Lookback limits: a breakpoint scans back at most 20 blocks for a prior write, and one developer saw misses at 23 blocks. Compaction run through a different system prompt misses the parent's prefix entirely, so the biggest transcript gets billed at full price. Scheduled runs start a brand-new session on each invocation, so the static prefix is rewritten from scratch; the post treats a near-zero cron hit rate as a consequence of architecture rather than a bug a longer TTL can fix.

The post also relays a reader's story: an agent whose logs looked flawless until a three-day-weekend invoice made no sense. The cause was a context-retrieval step pulling whole documents instead of chunks — 40,000 tokens per call, seven or eight calls per task, with every response technically successful.

Why it matters

Agent workloads are dominated by input tokens; the post measures one agent run at 3.5 million input tokens against 271,000 output. At that ratio, caching is where the budget is won or lost on the input side, and vendor-level dashboards average away precisely the callers doing the leaking. The recommendations are concrete: log hit rate per caller alongside the TTL tier, read the raw cache buckets from response usage rather than summaries, keep stable content ahead of cache_control with variable data such as date, working directory and branch behind it, apply 1-hour writes only on dispatch turns and shared static prefixes, and audit breakpoints that are written but never read. The lesson travels: OpenAI and Google handle caching differently, but wherever prefixes belong to conversations, per-caller accounting is what makes the invoice explainable.

  • #claude-code
  • #anthropic
  • #prompt-caching
  • #ai-agents
  • #cost-optimization

Related posts