deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

Explainer untangles what million-tokens-per-second claims mean for LLM context and output speed

A widely shared echohive.ai explainer separates 'one million tokens per second' into four distinct claims — context size, input speed, aggregate throughput and single-agent output — and shows why the differences matter.

Explainer untangles what million-tokens-per-second claims mean for LLM context and output speed

An interactive explainer published by echohive.ai and picked up on the Hacker News front page asks what would really change if an AI agent could write a million tokens per second. Its central point is that the phrase is ambiguous, and that most of the confusion around model speed claims comes from collapsing four distinct quantities into a single number.

One number, four meanings

According to the piece, "a million a second" can refer to context capacity — the text a model can hold in one request, including its reply, which is a size rather than a speed; input processing — how fast a model reads a prompt, which can be parallelized to a degree and differs from the output rate; aggregate throughput — what a serving system delivers in total, so 10,000 concurrent streams at 100 tokens per second each already add up to a million; and single-agent output — how fast one model actually writes, where standard generation is inherently serial because each token depends on the ones before it.

The explainer uses Anthropic's Claude Opus 5.5 overview to show how these ideas get conflated: it lists a one-million-token context window, a much smaller per-request output cap, and comparative latency described only as "moderate". A large context window says nothing about how quickly text is produced, and a million-token reply would not even fit in a single request.

Engineering raises the numbers but does not abolish the dependency chain, the piece argues. The 2023 vLLM paper describes how services reach large totals by batching many requests together, and speculative decoding lets several drafted tokens be accepted in one pass — yet neither removes situations where step two needs the result of step one. A footnote worth repeating: a token is a chunk of text, often part of a word, so token counts are not word counts.

What the imagined dial would buy

The page's slider runs from 100 to 1,000,000 tokens per second, and its worked examples — labelled imagined comparisons, not measurements — count generated output only, excluding reading, ranking, tool calls, tests and human review. At the slow end, 1,000 story endings of 1,000 tokens each take roughly 2 hours 46 minutes to write; at the imagined rate, about one second. Forty 20,000-token drafts of a lending app drop from more than two hours to 0.8 seconds. Critiquing 10,000 candidates at 300 tokens each falls from over eight hours to three seconds, and 10,000 one-thousand-token rehearsals of a difficult conversation go from more than a day to ten seconds.

Each thought experiment ends at a bottleneck the dial cannot touch. A thousand endings still need someone's taste to choose one. Forty app drafts can all pass tests that never check the seven-day due date. Ten thousand proposed experiments still wait on lab time measured in hours, days or a season, and ideas from one model may share a single blind spot. Rehearsing against the wrong person ten thousand times produces confidence, not accuracy.

Amdahl's law arrives

The sharpest section concerns fixed serial costs. Take the 800,000-token app batch and add one 60-second check after writing. At 100 tokens per second the job takes about 2 hours 14 minutes; at the imagined million, 60.8 seconds. Writing became 10,000 times faster, but the end-to-end job became roughly 132.57 times faster, with checking now 98.7 percent of the remaining time. Stretch the fixed step to 15 minutes and the gain shrinks to 9.88 times; stretch it to 24 hours and it nearly vanishes at 1.09 times. Whatever you do not accelerate becomes almost all of the remaining time, and the explainer cites OpenAI's latency guide to note that very large prompts, tool calls and network round trips add delays of their own in real requests.

Why it matters

Token-rate figures now appear in marketing, benchmarks and commentary with little consistency, and this explainer works as a checklist for reading them: is the claim about context size, input speed, batched throughput across many users, or the single stream you are actually waiting on? Only the last determines when your request finishes. The piece's larger conclusion is economic: when generation becomes nearly free, the scarce inputs become judgment and verification — taste, clear specifications, physical evidence, fidelity to real people — plus every serial step, human or mechanical, that sits outside the model.

  • #llm
  • #inference
  • #latency
  • #throughput
  • #ai

Related posts