deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Grok-4 drops to $2 per million input tokens while Qwen, Xiaomi and Z-AI ship new models

Grok-4's input price has been cut to $2 per million tokens, the only confirmed LLM price move this week, while fifteen new models from Qwen, Xiaomi and Z-AI landed in a single day.

Grok-4 drops to $2 per million input tokens while Qwen, Xiaomi and Z-AI ship new models

Grok-4 cut to $2 per million input tokens

The only confirmed LLM price change this week was a cut to Grok-4, which now costs $2 per million input tokens and $6 per million output tokens. The figures come from the weekly LLM Pricing Digest published on dev.to, which draws its data from LLM Price Watch, a tracker that checks pricing daily via OpenRouter across model families including Claude, GPT, Gemini, DeepSeek and Grok.

The cut matters because of where Grok-4 started. According to the digest, the model launched as a frontier-tier offering, a category where input pricing has reliably sat at $10 to $15 or more per million tokens. At $2 in and $6 out, it now overlaps with mid-tier pricing, which changes the calculation for production workloads rather than just experiments.

The practical comparison the digest draws is with teams routing complex reasoning tasks to GPT-4o or Claude Sonnet, typically paying somewhere between $5 and $15 per million input tokens depending on their mix. At $2 input, Grok-4 is now worth benchmarking against those incumbents. Output-heavy pipelines, such as long generations or document drafting, will notice the $6 output rate more than the input savings, but the digest still considers that competitive for a model at this capability level.

No movement from Claude, GPT or Gemini

The tracker recorded no price changes for Claude, GPT or Gemini models during the week. Grok-4's cut was the single confirmed movement, making it a quiet week for pricing but a busy one for new releases.

Fifteen new models in a single day

On September 30, fifteen new models appeared in the tracker, all arriving on the same day. The digest reads a same-day cluster like that as a coordinated launch rather than a staggered rollout.

Qwen shipped a set of qwen3.8 variants, including max-prime, omni-flash, max-0902, flash and a 27B model, plus a free tier of the 27B. The naming follows the family's established pattern: max signals higher capability, flash faster and cheaper, and omni multimodal. The free 27B tier is the standout for developers, since no-cost access to a model of that size suits prototyping and low-volume internal tools where spend is the binding constraint.

Xiaomi added three MiMo v2.6 models: pro-ultraspeed, flash and pro. The MiMo line is newer to the router ecosystem, and the three-way split mirrors the speed-versus-quality spectrum other vendors already cover. The digest had no pricing details to share beyond the tracker entries.

Z-AI, whose GLM series comes from Zhipu AI, was the most prolific of the three, with a GLM-5.3 family spanning prime, flashx, flash and base variants alongside batch-mode endpoints. Batch variants typically trade slower turnaround for lower per-token pricing, which matters for offline pipelines and other async workloads that do not need real-time responses.

What the digest suggests doing

For teams actively managing API costs, the digest's recommendations are:

  • Re-evaluate Grok-4 if it was written off earlier on price, since $2 input is a different conversation from where it started.
  • Treat the Qwen 27B free tier as a legitimate option for dev and test environments or low-stakes internal tools.
  • Look at GLM-5.3 batch endpoints for bulk processing, where batch pricing tends to undercut synchronous endpoints meaningfully.
  • Watch Xiaomi's MiMo line, but wait for more community benchmarks before routing production traffic to it.

Why it matters

A frontier-tier model at $2 per million input tokens lowers the cost floor for serious reasoning work, and it lands squarely in the $1 to $6 per million input band that the digest identifies as increasingly crowded. For developers with routing flexibility, every week like this widens the set of capable options available behind a single API and strengthens their leverage over cost. Grok-4's cut also applies indirect pressure on the incumbents: Claude, GPT and Gemini held their prices this week, but a competitor matching their tier at a fraction of the input cost is the kind of shift that could prompt a response. One caveat is that these numbers reflect a single tracker watching OpenRouter daily, so prices may move again quickly, and the digest itself points readers to the live tracker for the latest data.

  • #llm
  • #api-pricing
  • #grok-4
  • #qwen
  • #openrouter

Related posts