deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Claude Sonnet 5.5 arrives with a caveat: some prompts may be silently rerouted to Sonnet 5

Anthropic released Claude Sonnet 5.5 with stronger agentic coding benchmarks than Opus 5.5 at unchanged Sonnet 5 pricing, but higher-risk requests can be quietly rerouted to the older model.

Claude Sonnet 5.5 arrives with a caveat: some prompts may be silently rerouted to Sonnet 5

Anthropic released Claude Sonnet 5.5 on September 28, six days after Opus 5.5, and the rollout carries an unusual admission: under certain conditions, a request sent to the new model may be answered by the older Sonnet 5. According to a dev.to write-up summarising the launch, The New Stack reported that in situations Anthropic describes as "higher-risk", prompts aimed at Sonnet 5.5 can be rerouted to Sonnet 5 rather than refused or flagged, and the criteria that trigger the swap have not been published.

Stronger numbers at unchanged prices

Sonnet 5.5 keeps the previous model's pricing — $2 per million input tokens and $10 per million output tokens — while improving results across the benchmark figures relayed in the dev.to post. On Terminal-Bench 4.0, which measures agentic coding where a model plans, edits files and iterates, Sonnet 5.5 scored 70.6%, ahead of the flagship Opus 5.5 at 66.4% and far beyond Sonnet 5's 10.3%. On GDPval-AA v2.1, a general knowledge-work evaluation, it reached 1844 — two points shy of Opus 5.5's 1846 and roughly 400 points above Sonnet 5's 1449. On Chartography, a visual chart recognition test, it hit 61.6%, trailing Opus 5.5's 64.4% but up sharply from Sonnet 5's 15.6%. Anthropic also claims output is at least 30% faster and per-task cost up to 30% lower, which the company attributes to the model needing fewer tokens and fewer tool calls to finish a job.

Taken together, the cheaper model now leads the flagship on the benchmark most relevant to agent-based coding, which the dev.to author argues makes it the sensible default for cost-sensitive agentic pipelines rather than Opus 5.5.

The silent rerouting caveat

The model ships with safeguards of a kind Anthropic until now kept for its most capable models: classifiers designed to detect and block misuse patterns such as attempts to extract the model's reasoning. The downside, as reported by The New Stack, is that these safeguards can redirect a request. When Anthropic judges a prompt to be higher-risk, the response may come from Sonnet 5 — a model with different capabilities — even though the developer selected Sonnet 5.5. Nothing is refused and nothing is flagged, and because the trigger conditions are undocumented, there is no way to design around the behaviour, only to be surprised by it.

That matters most wherever consistent model identity is assumed: running evaluations, benchmarking an internal pipeline, debugging a regression or reproducing a reported bug. A swap the caller cannot see turns "it worked yesterday" from a debugging clue into a matter of chance.

Heavy early usage

The launch became the top story on Hacker News, drawing 775 points and more than 500 comments. Commenters cited by the dev.to post pointed to token efficiency rather than hype: one said they had largely stopped using Sonnet after Opus 5.5 arrived but returned because the new release consumes far fewer tokens, and another reported their Pro-plan quota lasting a full day for the first time in months. A third exhausted a week's usage allotment and its reset within hours — a sign, the post suggests, that people are pointing real workloads at the model rather than testing it.

Why it matters

The benchmark gains are straightforwardly good news for anyone running coding agents at volume: near-Opus performance on agentic tasks at a fraction of the price. The routing admission is the part with longer consequences. APIs are built on the premise that naming a model pins down its behaviour, and evals, regression suites and compliance logs all depend on that premise holding. If a provider can substitute a different model mid-request without notice, verification shifts to the caller. The practical fix, as the dev.to post puts it, is to log the model field returned in every response rather than trusting the model that was requested — currently the only reliable way to catch a swap. Until Anthropic documents what counts as higher-risk, Sonnet 5.5 is both the better economic choice for agentic work and a model whose identity cannot be fully guaranteed.

  • #anthropic
  • #claude
  • #llm
  • #api
  • #model-routing

Related posts