· via dev.to (home feed)
Anthropic's Opus 5.5 guide says drop 'think step by step' and watch for silent downgrades
Anthropic's Opus 5.5 prompting guide reportedly says 'think carefully' prompts now only add latency, long runs need defined finish lines, and flagged messages can silently fall back to an older model.

The model manages its own reasoning now
Anthropic's engineering blog recently published a guide by Addy Osmani on getting the most from Opus 5.5 in Claude and Claude Code, and a dev.to write-up summarizing it notes the piece reached the Hacker News front page. Its central argument is uncomfortable for anyone who spent years collecting prompt tricks: Opus 5.5 deliberates before every reply and judges for itself how much thought a task deserves, so nudging it toward reasoning is redundant. In Anthropic's own chat-product testing, cited by dev.to, removing a "think carefully" line let responses start sooner with no clear drop in quality. The instruction now only adds delay, and in saved instructions it adds that delay to every message you send.
One message, a finish line, a stop condition
Osmani's recommended structure hands over the complete task up front in three parts: the task itself, such as migrating payment endpoints to a new client; a definition of done, such as every endpoint migrated, the old client deleted and the test suite passing; and the single condition that should interrupt the run, such as a test failing for reasons the model cannot explain. Everything else should be a status note, not a question. The rationale, per dev.to, is that early testers left Opus 5.5 working through large repositories for hours with little oversight, and its strongest gains over Opus 5 appeared exactly there.
For design work, the guide favors specific negatives over vague positives. Telling the model to avoid a generic look mostly swaps one generic result for another, while a concrete ban list — no off-white backgrounds, no italic accent words, no pill-shaped buttons — performs better, and anything you dislike gets added to the list and rerun.
CLAUDE.md becomes run control
The guide repurposes CLAUDE.md from a style reference into a steering mechanism for agent sessions. The suggested rules tell the model to keep going when a step needs no human input, folding status notes into its next action, and to halt only when it truly cannot proceed or before anything destructive — deleting data, force-pushing, or changing anything outside the repository. The destructive-command carve-out matters: fewer stops means keeping permission prompts enabled for risky operations remains your safety net.
One behavior the guide flags is the model pausing mid-task to report, naming the next step without taking it. The keep-going rule is the systematic fix, and replying "continue" handles one-off cases.
For audits and migrations spread across a codebase, Osmani suggests one subagent per service, with a twist: the model itself should verify each subagent's evidence before accepting it, closing with a single table of service, affected status and evidence. Because Claude Code summarizes older turns when long runs fill the context window, the guide also recommends keeping the task checklist in a file such as TASKS.md, which survives that summarization.
Review the output deliberately
When a run ends, read what is blocked on you first — an open decision or a change awaiting approval — before the rest of the summary, and standardize the ending by requiring three closing headings: blocked on me, changed, and found. Other habits include running a model review before the human one, with one early tester reporting that Opus 5.5 at its lowest effort caught more bugs than Opus 5 at high effort, with fewer false alarms; asking research tasks to mark unconfirmed claims and state where they looked; and attaching charts and screenshots instead of retyping numbers, since Opus 5.5 reportedly reads images, including spatial relationships like which boxes an arrow connects, more accurately than its predecessor.
Silent downgrades are the overlooked risk
The section dev.to calls underdiscussed concerns model routing. Opus 5.5 is the first Opus model to launch with Fable-level bio and cyber safeguards, and in Claude apps and Claude Code most flagged messages are not refused — they are quietly rerouted to an older model while the session continues. The check covers the entire conversation, including files and search results, so earlier content can trigger it. In Claude Code, the /model command switches back or pressing Escape twice lets you edit and retry, and a /config setting forces the model to ask before switching; in the apps, reselect the model or start a new chat, with the same toggle under Settings. Notably, asking the model to reproduce its internal reasoning is itself a flag category and can be declined outright, so request a brief rationale instead.
Fast mode
For interactive back-and-forth, Claude Code offers /fast: the same model with text arriving sooner. Per the guide it is a research preview, requires extra usage to be enabled and costs more per token, which makes it a poor fit for long autonomous runs where nobody is reading each reply.
Why it matters
Vendor prompting guidance directly shapes how thousands of developers write system prompts and saved instructions. If Anthropic's testing holds up, legacy reasoning prompts now buy latency and nothing else, and the craft shifts from eliciting thought toward defining completion criteria, stop conditions and guardrails. The silent downgrade mechanism deserves equal attention: anyone doing security review or adjacent work could be receiving older-model output for hours without noticing, and the ask-first toggle is cheap insurance against exactly that failure.
- #anthropic
- #claude
- #prompting
- #ai-agents
- #llm