deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Grok 4.7 doubles cost per task despite unchanged $2/$6 token pricing

Artificial Analysis measurements show Grok 4.7 emits about 2.5x more tokens than Grok 4.6 at identical per-token prices, lifting cost per task from $1.86 to $3.74.

Grok 4.7 doubles cost per task despite unchanged $2/$6 token pricing

Unchanged price list, doubled bill

Grok 4.7 launched at exactly the same per-token rates as Grok 4.6 — $2 per million input tokens and $6 per million output tokens — yet a typical task now costs roughly twice as much. According to figures from Artificial Analysis reported by dev.to, Grok 4.7 generated about 240 million tokens while running the Intelligence Index benchmark suite, compared with 94 million for Grok 4.6 and a median of 88 million across models. Average cost per task jumped from $1.86 to $3.74.

The arithmetic is simple: what you pay for a task is the price per token multiplied by the number of tokens the model decides to produce, and reasoning models set that second number themselves. Grok 4.7 wrote roughly 2.5 times as much as its predecessor to gain two points on the index, which moved from 44 to 46.

One caveat flagged by dev.to: the headline comparison runs Grok 4.7 at its "xhigh" reasoning-effort setting against Grok 4.6 at "high". At matched effort, the Artificial Analysis article still reports about 81,000 output tokens per task for Grok 4.7 versus 38,000 for Grok 4.6, and roughly triple the figure cited for GPT-6 Astra. Either way, token counts more than double, and each Intelligence Index task took about 7.1 minutes to complete.

The speed claim versus the stopwatch

xAI's launch post headlines the model as twice as fast at half the price of comparable models, then states a few paragraphs later that it is served at the same price and speed as Grok 4.6. As dev.to points out, both statements can only hold if the speed claim refers to other labs' flagships — the post's own pricing table lists competing models at $4/$20 and $10/$50 per million tokens — rather than to the previous Grok, which cost the same.

The independent numbers point the other way. Artificial Analysis clocked output at about 57 tokens per second, below Grok 4.6's 66 and a field median of 79, and its summary called the model notably slow. The launch post does describe a new, larger base model and a longer reinforcement-learning run aimed at problems that take hours. A widely discussed comment in a 609-point Hacker News thread claimed the model carries 40% more weights than Grok 4.6 at the same price, but dev.to notes that xAI published no such figure, so the number remains a commenter's estimate.

The verbosity bought real gains

The extra tokens are not pure waste. Artificial Analysis reported improvements concentrated exactly where xAI said it trained the model: long, multi-step work. AA-Briefcase, which covers lengthy office tasks, rose 111 Elo to 1657; GDPval-AA rose 90 points to 1695; and the Coding Agent Index, run in xAI's own Grok Build harness, climbed nine points to 56, ranking fourth behind Claude Fable 5.1, GPT-6 Astra and Claude Opus 5. The hallucination rate dropped from 34% to 29%, and the model sits 21st of 211 on the overall Intelligence Index.

Lazy, loopy, or both

Developer reports from the Hacker News thread, as relayed by dev.to, pull in contradictory directions. Some users say the model ends tasks almost immediately and declares the work done; others describe it looping through dozens of enumerated fixes in thinking mode while ignoring AGENTS.md instructions and corrupting plan files. A minority called it a focused model that stays on course. There is no contradiction, dev.to argues: a model can spend 81,000 tokens on an elaborate plan and still stop before the work is finished — leaving the user to fund the planning and then do the job.

Why it matters

For anyone budgeting LLM usage, this launch is a concrete demonstration that a provider's price table no longer predicts the bill. The provider fixes the rate; the model chooses the volume. The practical advice from dev.to is worth repeating: compare cost per task rather than price per token, using published measurements or your own logs; check which reasoning-effort level your SDK or provider defaults to before comparing models; and set per-task token or spend caps in your agent harness before switching. The episode is also a case for independent measurement — Artificial Analysis both confirmed xAI's claimed gains on long-horizon tasks and contradicted its speed claim, and a buyer who read only the launch post would have budgeted about half of what the model actually costs.

  • #llm
  • #grok
  • #pricing
  • #benchmarks
  • #artificial-analysis

Related posts