deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Anthropic launches Claude Sonnet 5.5 with Terminal-Bench lead at $2 per million tokens

Anthropic has released Claude Sonnet 5.5, a mid-tier model that a dev.to report says beats larger rivals on Terminal-Bench 4.0 while costing $2 per million input tokens.

Anthropic launches Claude Sonnet 5.5 with Terminal-Bench lead at $2 per million tokens

What launched

Anthropic has officially released Claude Sonnet 5.5, the mid-tier model in its current lineup, according to a dev.to post published on 11 October. The model sits between the lightweight Haiku 5.5 and the deep-reasoning Opus 5.5, yet in terminal-based software work it reportedly beat not only its larger sibling but also more expensive frontier rivals such as OpenAI's GPT-6 Astra and Google DeepMind's Gemini 4 Argon.

The headline result is a 70.6% resolution rate on Terminal-Bench 4.0, a benchmark that asks a model to diagnose a bug in a real repository, install missing dependencies, run tests, interpret stack traces, apply a fix and confirm the build without human help. According to the dev.to write-up, that is more than four points clear of Opus 5.5 at 66.4% and almost nine ahead of GPT-6 Astra at 61.8%, with Gemini 4 Argon and the open-weight DeepSeek 4.1 trailing at 57.4% and 55.2%.

Strength beyond the command line

The post also highlights multimodal gains. Sonnet 5.5 scored 61.6% on the Chartography visual-reasoning benchmark without external OCR tools or helper scripts, reportedly handling dense architectural diagrams and data-heavy vector graphics. On OSWorld 2.1, which measures desktop GUI automation, the model reached 80.1%, described in the post as among the strongest results available for operating-system task automation. As a demonstration of sequential visual reasoning, the model was set the challenge of playing the Game Boy classic Pokémon Red using only raw screenshots, reading the interface, navigating battle menus and planning routes across the map entirely from what it saw.

Where Opus still leads

Opus 5.5 keeps narrow leads on deep synthesis and IDE-heavy work. The comparison table in the post shows Opus ahead on CursorBench 4.0 (57.8% against Sonnet's 55.5%), FrontierCode 1.1 in its Xhigh configuration (54.3% versus 52.1%), GDPval-AA v2.1 (1,846 versus 1,844 points) and AA-Briefcase (1,822 versus 1,811). On the knowledge-heavy benchmarks the margin is effectively a rounding error, suggesting the mid-tier model now sits within touching distance of the flagship on analytical breadth while costing half as much for input tokens.

Speed, context and price

For daily developer tooling such as Claude Code CLI, Cursor and Windsurf, the post reports output tokens arriving more than 30% faster than the previous mid-generation model. The release pairs that throughput with a 128,000-token output ceiling and a one-million-token context window. A refined Preserved Thinking mechanism keeps earlier reasoning resident in context memory with active caching across successive tool calls, so the model does not have to rebuild its assumptions from scratch after every shell command it executes.

Pricing is central to the story. The post lists Sonnet 5.5 at $2.00 per million input tokens, against $4.00 for Opus 5.5, $12.00 for GPT-6 Astra and $10.00 for Gemini 4 Argon; only DeepSeek 4.1 is cheaper at $0.55. The model is also catalogued on Models.dev, where the post says it is being tracked in real time.

Why it matters

If the reported numbers hold up, they mark a shift in what frontier means in 2026: the decisive competition is no longer peak reasoning on one-off problems but the cost and latency of agents running hundreds of iterations per hour. A mid-tier model that beats flagships on autonomous terminal work at a sixth of the input price of some rivals changes the economics of continuous background agents in CI pipelines and production engineering loops. One caveat applies, however: every figure in this story comes from a single community post on dev.to, and no independent benchmark run or Anthropic announcement document was available among the sources to confirm them.

  • #anthropic
  • #claude
  • #llm
  • #benchmarks
  • #ai-agents

Related posts