· via dev.to (home feed)
Grok-4 drops to $2 per million input tokens while Qwen, Xiaomi and Z-AI ship new models
Grok-4's input price has been cut to $2 per million tokens, the only confirmed LLM price move this week, while fifteen new models from Qwen, Xiaomi and Z-AI landed in a single day.

Grok-4 cut to $2 per million input tokens
The only confirmed LLM price change this week was a cut to Grok-4, which now costs $2 per million input tokens and $6 per million output tokens. The figures come from the weekly LLM Pricing Digest published on dev.to, which draws its data from LLM Price Watch, a tracker that checks pricing daily via OpenRouter across model families including Claude, GPT, Gemini, DeepSeek and Grok.
The cut matters because of where Grok-4 started. According to the digest, the model launched as a frontier-tier offering, a category where input pricing has reliably sat at $10 to $15 or more per million tokens. At $2 in and $6 out, it now overlaps with mid-tier pricing, which changes the calculation for production workloads rather than just experiments.
The practical comparison the digest draws is with teams routing complex reasoning tasks to GPT-4o or Claude Sonnet, typically paying somewhere between $5 and $15 per million input tokens depending on their mix. At $2 input, Grok-4 is now worth benchmarking against those incumbents. Output-heavy pipelines, such as long generations or document drafting, will notice the $6 output rate more than the input savings, but the digest still considers that competitive for a model at this capability level.
No movement from Claude, GPT or Gemini
The tracker recorded no price changes for Claude, GPT or Gemini models during the week. Grok-4's cut was the single confirmed movement, making it a quiet week for pricing but a busy one for new releases.
Fifteen new models in a single day
On September 30, fifteen new models appeared in the tracker, all arriving on the same day. The digest reads a same-day cluster like that as a coordinated launch rather than a staggered rollout.
Qwen shipped a set of qwen3.8 variants, including max-prime, omni-flash, max-0902, flash and a 27B model, plus a free tier of the 27B. The naming follows the family's established pattern: max signals higher capability, flash faster and cheaper, and omni multimodal. The free 27B tier is the standout for developers, since no-cost access to a model of that size suits prototyping and low-volume internal tools where spend is the binding constraint.
Xiaomi added three MiMo v2.6 models: pro-ultraspeed, flash and pro. The MiMo line is newer to the router ecosystem, and the three-way split mirrors the speed-versus-quality spectrum other vendors already cover. The digest had no pricing details to share beyond the tracker entries.
Z-AI, whose GLM series comes from Zhipu AI, was the most prolific of the three, with a GLM-5.3 family spanning prime, flashx, flash and base variants alongside batch-mode endpoints. Batch variants typically trade slower turnaround for lower per-token pricing, which matters for offline pipelines and other async workloads that do not need real-time responses.
What the digest suggests doing
For teams actively managing API costs, the digest's recommendations are:
- Re-evaluate Grok-4 if it was written off earlier on price, since $2 input is a different conversation from where it started.
- Treat the Qwen 27B free tier as a legitimate option for dev and test environments or low-stakes internal tools.
- Look at GLM-5.3 batch endpoints for bulk processing, where batch pricing tends to undercut synchronous endpoints meaningfully.
- Watch Xiaomi's MiMo line, but wait for more community benchmarks before routing production traffic to it.
Why it matters
A frontier-tier model at $2 per million input tokens lowers the cost floor for serious reasoning work, and it lands squarely in the $1 to $6 per million input band that the digest identifies as increasingly crowded. For developers with routing flexibility, every week like this widens the set of capable options available behind a single API and strengthens their leverage over cost. Grok-4's cut also applies indirect pressure on the incumbents: Claude, GPT and Gemini held their prices this week, but a competitor matching their tier at a fraction of the input cost is the kind of shift that could prompt a response. One caveat is that these numbers reflect a single tracker watching OpenRouter daily, so prices may move again quickly, and the digest itself points readers to the live tracker for the latest data.
- #llm
- #api-pricing
- #grok-4
- #qwen
- #openrouter