deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

xAI ships Grok 4.7, its strongest coding model, at unchanged $2/$6 per million token pricing

xAI says Grok 4.7 is its most capable model for coding and knowledge work, delivered at the same price and speed as Grok 4.6, with big gains on long-running agentic benchmarks.

xAI ships Grok 4.7, its strongest coding model, at unchanged $2/$6 per million token pricing

xAI has released Grok 4.7, a model the company calls its strongest yet for coding and knowledge work. The headline detail is economic rather than technical: according to xAI's announcement, Grok 4.7 is served at the same price and speed as Grok 4.6, meaning $2 per million input tokens and $6 per million output tokens, despite running on a new, larger base model.

What changed under the hood

Grok 4.7 was built on a bigger foundation model than its predecessor and put through a longer reinforcement learning run on a harder mixture of tasks. xAI says that training mix was deliberately weighted toward problems that take many hours to finish, which shows up in the model's stronger results on long-running work.

The company also highlights two other improvements: the model checks its own output more reliably, and it manages longer context windows better. It was additionally trained to natively understand the Grok Bot harness, which xAI credits for better conversational behaviour and general knowledge performance.

Benchmark results, as reported by xAI

On CursorBench 4.0, a benchmark focused on extended coding sessions, Grok 4.7 scores 46.3%, up from Grok 4.6's 40.4% and ahead of GPT-5.6 Sol Max at 41.7%. Fable 5.1 Max still leads that comparison at 51.8%.

The gains are largest on tasks that reward persistence. On Terminal-Bench 4.0, which measures multi-hour terminal work, Grok 4.7 more than doubles Grok 4.6's score, reaching 38.0% versus 20.3%. On DeepSWE v1.1, it posts 71.0% in a high-effort configuration, narrowly behind GPT-5.6 Sol's 72.7% and just ahead of Fable 5.1 at 70.0%.

The model also targets professional knowledge work. On EEBench, an electrical engineering evaluation, it scores 64.0% against 53.0% for Grok 4.6. On the Harvey Legal Agent Benchmark it reaches 19.6%, far ahead of the other models listed, which score in the single digits. On GDPval, where models handle tasks normally done by lawyers, nurses and financial analysts, Grok 4.7 earns an Elo of 1695, ahead of Grok 4.6 and GPT-6 Astra but behind Fable 5.1 at 1735. Its 56.7% on HealthBench Professional trails both GPT-5.6 Sol and Fable 5.1.

Safety and cybersecurity claims

xAI says Grok 4.7 ships with an entirely new safeguard stack and is the strongest model it has tested on refusals and jailbreak resistance. On LatchBio's biosafety benchmark it tops the field at 62.4%, and on xAI's own HackerBench v0.3 it lets only 3.3% of risky dual-use prompts through while rarely blocking legitimate security work. The company has also begun offering select cybersecurity partners invite-only access to the model's red-team capabilities for defensive research.

Pricing and availability

Grok 4.7 is available now in Cursor and Grok Build, and through the Grok API, third-party coding harnesses, model routers and cloud platforms. A fast variant doubles output speed at double the price. All benchmark and pricing figures above come from xAI's own announcement.

Why it matters

The competitive pressure here is price-performance. xAI's comparison table shows rival frontier models costing several times more per token: GPT-5.6 Sol Max at $4 per million input and $20 per million output, and Fable 5.1 Max at $10 and $50. If xAI's numbers hold up, Grok 4.7 delivers frontier-adjacent coding ability at a fraction of that cost, which matters for anyone running high-volume agentic workloads where token bills dominate.

The gains are also concentrated exactly where the industry is heading: multi-hour, self-directed tasks in terminals, IDEs and office-work harnesses rather than single-turn chat. That said, every score cited is vendor-reported and measured on benchmarks xAI selected, so independent evaluation will determine whether the price-performance claim survives contact with real-world usage.

  • #xia
  • #grok
  • #large-language-models
  • #ai-coding
  • #benchmarks

Related posts