deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Claude Sonnet 5.5 nears Opus 5.5 on agentic coding tests at half the price

Anthropic's new Claude Sonnet 5.5 beats Opus 5.5 on Terminal-Bench 4.0 while keeping Sonnet 5 pricing, adding faster output and a November retirement date for Sonnet 4.5.

Claude Sonnet 5.5 nears Opus 5.5 on agentic coding tests at half the price

What Anthropic shipped

Anthropic released Claude Sonnet 5.5 on September 28, 2026, a mid-tier model that lands close to its flagship Opus 5.5 on agentic coding evaluations while costing half as much per token. According to Anthropic's announcement, the new model generates output more than 30% faster than Sonnet 5 and needs fewer tokens for the same job, which the company says can reduce the cost of a task by up to 30%.

Pricing is unchanged from Sonnet 5: $2 per million input tokens and $10 per million output tokens, with cache reads at $0.20 and cache writes at $2.50 per million. Opus 5.5, released a week earlier on September 22, is exactly double at $4 and $20 per million. Developers call the model with the ID claude-sonnet-5-5, and it is available on the Claude Platform, Amazon Web Services, Google Cloud, Microsoft Azure and in the Claude apps. Per MarkTechPost, it carries a 1 million token context window, 128,000 maximum output tokens and a June 2026 knowledge cutoff.

Benchmark results

Anthropic published results comparing Sonnet 5.5 against both Sonnet 5 and Opus 5.5. The headline number is Terminal-Bench 4.0, a test of agents working at the command line, where Sonnet 5.5 scores 70.6% and edges past Opus 5.5 at 66.4%.

On most other benchmarks Opus 5.5 keeps a narrow lead: CursorBench 4.0 (55.5% versus 57.8%), FrontierCode 1.1 Main (46.2% versus 54.4%), OSWorld 2.1 (80.1% versus 81.8%), Humanity's Last Exam (64.5% versus 67.7%), Chartography (61.6% versus 64.4%) and GDPval-AA v2.1 (1844 versus 1846).

The generational jump over Sonnet 5 is sharpest on agent tasks. Sonnet 5 managed only 10.3% on Terminal-Bench 4.0 and 15.6% on Chartography, making command-line agent work the largest improvement in Anthropic's table.

How customers are using it

Anthropic positions the two models as complements: Sonnet 5.5 for well-scoped everyday work such as bug fixes, document generation and fast iteration, and Opus 5.5 for open-ended problems requiring sustained reasoning.

Early adopters report measurable gains. Slack principal engineer Curtis Allen said that without changing prompts, Sonnet 5.5 performed better on nearly all of the company's evals while using about 14% fewer output tokens. App builder Base44 measured 3.6 iterations per build with Sonnet 5.5 against 7.7 for Opus 5, according to Unite.AI. Atlassian says its Rovo Agents can run up to 30% faster than with Sonnet 5.

Safety guardrails and a fallback

Anthropic says Sonnet 5.5 is the first Sonnet model with safety classifiers that block attempts to extract its reasoning. For higher-risk cybersecurity tasks, requests will visibly fall back to Sonnet 5, and security researchers can apply for broader access through Anthropic's Cyber Verification Program.

Sonnet 4.5 retirement clock

The release also starts a countdown for older models. Anthropic's deprecation page says Claude Sonnet 4.5 (claude-sonnet-4-5-20250929) was deprecated on September 30, 2026 and retires on November 30, 2026, with claude-sonnet-5-5 named as the replacement. Amazon Bedrock and Google Cloud set their own retirement dates.

Why it matters

A mid-tier model that matches near-flagship performance at half the token price resets the economics of AI-assisted coding. Teams that defaulted to Opus for agent workloads now have a credible cheaper route: Anthropic's own positioning suggests routing well-defined coding, bug fixing and document tasks to Sonnet 5.5 while reserving Opus 5.5 for problems that need deeper judgment. The 30% faster output and reduced token use compound the price difference on high-volume pipelines.

There are two operational catches. First, anyone still calling Sonnet 4.5 has until November 30 on Anthropic's platform, and separate deadlines apply on Bedrock and Google Cloud. Second, the visible fallback to Sonnet 5 on higher-risk security requests means responses can silently come from a different model, so logging which model answered each request is prudent. As always with benchmark-driven releases, running your own evals against real prompts before switching remains the sensible default.

  • #anthropic
  • #claude
  • #ai-models
  • #agentic-coding
  • #llm-pricing

Related posts