· via dev.to (home feed)
Hidden LLM bills: context bloat from MCP tool schemas and skills, plus shadow eval spend
Two dev.to analyses tally what LLM agents cost before and after they run: context consumed by MCP schemas and skill files, and the unbudgeted spend on evaluating whether outputs actually work.

Two token bills before the agent answers
The first analysis, published on dev.to, measures what an agent already carries in its context window before a user types anything. The author tokenized a reference setup with tiktoken, OpenAI's tokenizer, using the cl100k_base encoding, and found two separate charges.
The first is MCP tool schemas. Every connected Model Context Protocol server ships the full JSON description of every tool it exposes, and the manifest is resent on every turn. A catalog of 255 tools rendered as complete schemas came to 71,929 tokens — a cost paid once per question in a session, not once per session.
The second, larger charge is skills. A skill is an entire document — instructions, examples, a workflow — rather than a name plus a schema, and the author's skills directory held 418 SKILL.md files totalling 1,109,242 tokens when concatenated. That cannot fit in a context window, so most setups load a guessed subset and pay a few thousand tokens per skill on every turn. The author argues the cost goes unnoticed because it is spread across many small, reasonable-looking loading decisions.
The proposed fix is the same for both: deciding whether to use a tool or skill only requires its name, so expose an index of names alone and fetch full bodies on demand. The 255-tool catalog shrinks to 581 tokens as a name index, a 99.2 percent reduction, and a single resident pointer of 39 tokens replaces the skills catalog, with each lookup costing 501 tokens. The author ships the pattern as mcptoon, an open-source, zero-dependency Python CLI. The post also cites a Firecrawl benchmark that measured the same task at 1,365 tokens through a CLI versus 44,026 through MCP, a 32x gap, and notes that an Anthropic engineering write-up has already made the case on the schema side.
Both figures come from one machine, and the author is explicit that they indicate scale rather than a constant: a small setup with a couple of servers and a dozen skills should measure first and skip the extra layer.
The shadow bill: paying to prove it works
The second analysis, from The Agent Loop on dev.to, looks at the cost that arrives after the run: evaluation. Its framing is that LLM-as-judge scoring is a second inference workload with its own multipliers. It quotes Arize's formula: production evaluation cost equals traffic volume times sampling rate times evaluation surfaces times evaluator cost, plus human review plus retention.
How large can that get? The post cites a Monte Carlo report that one data leader described evaluation costs at 10 times the baseline agent workload — and flags this explicitly as a single anecdote, not a benchmark.
The defensible multipliers come from the literature. The τ-bench paper reports that its best gpt-4o function-calling agent exceeds 60 percent average task success while pass^8 drops below 25 percent, so demonstrating reliability demands rollouts; the paper prices one trial per task at roughly $200, with $0.38 for the agent and $0.23 for the simulated user per task. The post adds a caution: pass^8 is an eight-fold rollout opportunity, not necessarily an eight-fold dollar bill, since caching, sampling and trajectory lengths sit in between. Separately, a judge-tuning paper (arXiv 2501.17178) estimated about $2,000 to search 4,480 judge configurations versus roughly $2 million for full Alpaca-Eval-style evaluation, where one annotation costs about $24. The Agent Loop states it found no primary source for a universal "evals cost 5–30x a run" constant.
The recommended ladder keeps the shadow bill small: deterministic checks first, such as exit codes, schema validation and comparing final database state, which is how τ-bench itself scores; sample production traffic instead of scoring everything; escalate on uncertainty and consequence; use a small, cheap judge model for the easy stratum; keep humans for calibration and the high-consequence tail.
The coda is uncomfortable: cheap verification is not automatically correct verification. The post recounts Anthropic's documentation of Opus 4.5 scoring 42 percent on CORE-Bench until a researcher found rigid grading that penalized an answer of 96.12 when roughly 96.125 was expected, after which the score jumped to 95 percent.
Why it matters
Both analyses converge on one conclusion: the visible inference bill is not the whole bill. Context overhead is a per-turn multiplier that scales with how many tools and skills you have wired up, and evaluation is a largely unbudgeted workload that scales with how much you verify. Neither the million-token skill catalog nor the 10x eval anecdote is universal. The shared prescription is to measure your own numbers first — a tokenizer count on one side, a cost formula on the other — and to treat any single quoted ratio with suspicion until you have run it on your own workload.
- #llm-agents
- #mcp
- #ai-evaluation
- #token-costs
- #context-window