deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

Anthropic's Claude Opus 5.5 prompting guide maps effort, thinking and agent changes

Anthropic's new prompting guide details how Claude Opus 5.5 differs from Opus 5 — faster output, always-on thinking — and how to recalibrate effort, max_tokens and agent harnesses.

Anthropic's Claude Opus 5.5 prompting guide maps effort, thinking and agent changes

Anthropic has published a prompting guide for Claude Opus 5.5 that lays out how the model behaves differently from Claude Opus 5 and what developers should change in their prompts and harnesses as a result. The page, part of Anthropic's platform documentation, drew attention on Hacker News's front page — a sign of how much migration work it saves for anyone already running Opus in production.

What changes from Opus 5

According to the guide, Claude Opus 5.5 emits output tokens more than 30 percent faster than Claude Opus 5 and typically finishes the same task with fewer of them. Prompts written for Opus 5 should keep working unmodified, and Anthropic frames the older model's prompting patterns as a still-reasonable starting point.

On capability, Anthropic's testing found the model strongest at multistep work in a real repository — carrying a change through a large codebase until its tests pass. At its default medium effort it matched or beat Opus 5 running at high effort on such tasks, in fewer steps and with fewer tokens. It also sustains long autonomous runs better, including multi-hour audits and migrations executed end to end with parallel subagents and little oversight, and early testers reported sharper code review, with more real bugs caught and fewer false alarms.

In knowledge work, Anthropic says the model is much less likely to state a wrong figure or cite the wrong source, and it catches easy-to-miss details in large inputs — a date in a long planning thread falling on the wrong weekday, or a slide-deck chart that contradicts the underlying figures. Visual handling improved as well: even at its lowest effort setting the model read values off dense charts more accurately than Opus 5 did at its highest, using a small fraction of the output tokens, and it handles position-dependent meaning — which boxes an arrow connects, or when a meeting starts and ends in a calendar screenshot — more reliably. At default effort it matched the computer-use success rate Opus 5 reached only at a much higher setting.

Effort becomes the main dial

Thinking is always on in Opus 5.5, which makes the effort setting the primary lever for trading intelligence against latency and cost. The default drops from high on Opus 5 to medium on 5.5, and Anthropic is explicit that effort names are not comparable across models: medium on 5.5 matched or exceeded high on 5 in its coding and knowledge-work evaluations, and low came close on several coding evals at much lower cost. The advice is to set effort explicitly and test several levels against your own evals rather than inherit the old value.

At any given level the new model thinks more per turn, especially at xhigh and max, so an unchanged setting means longer turns and more tokens. Three adjustments follow: raise max_tokens, since thinking counts toward the limit even when the thinking content is not returned (Anthropic reports 128,000, the model's maximum, working well for long agentic turns); reserve xhigh and max for work where you have measured a quality gain; and to reduce thinking, lower effort rather than add prompt instructions, which Anthropic describes as less reliable. One operational wrinkle: changing the top-level effort between requests invalidates the prompt cache, so the guide points to a beta per-message effort change for individual turns.

Migrating integrations that disabled thinking

Opus 5.5 no longer accepts thinking disabled. For integrations that ran Opus 5 that way, the guide prescribes: start at low effort and measure latency and quality on your own traffic, moving up if quality drops; optionally add a system-prompt line such as "Answer directly without deliberating." if time to first token still matters, while re-measuring quality since less thinking can lower it; strip instructions that asked the model to write its reasoning into the response — Anthropic warns such prompts can be declined under a reasoning_extraction refusal category — and read summarized thinking blocks instead; re-test the old thinking-disabled mitigations, since the artifacts they addressed may not appear when thinking is always on; and parse responses by content block type rather than assuming the first block is text.

The rest of the map

The guide is organized as symptom-to-fix. A stop_reason of "refusal" routes to a section on safeguard refusals; silent long agentic turns route to user-facing progress updates; agents that miss information a task never pointed at route to exploring context in multi-app workflows; slow chat replies caused by lengthy up-front thinking route to thinking instructions in chat system prompts; instructions obeyed from user-pasted text route to marking pasted content; generic frontend output routes to design defaults; and misses on dense charts and screenshots route to tooling for complex visual inputs. A section on time signals targets multiagent harnesses that need to finish sooner.

Why it matters

Migration failures usually live in the harness, not the model. By publishing measured guidance on which old defaults break, which cheaper settings now suffice, and where cache invalidation and refusal handling bite, Anthropic turns an opaque model swap into a checklist. For cost-sensitive deployments, the claim that low effort on 5.5 approaches high effort on 5 at much lower cost may matter more than any capability jump.

  • #anthropic
  • #claude
  • #prompt-engineering
  • #llm
  • #ai-agents

Related posts