· via Hacker News – Front Page (native)
Anthropic's Opus 5.5 guide: longer autonomous runs mean new prompting habits
Anthropic's Opus 5.5 guide says the model thinks before every reply and sustains long coding runs, so 'think carefully' prompts should go and CLAUDE.md stop rules should take their place.

Anthropic's guide to a changed model
Anthropic has published a usage guide for Opus 5.5, its newest Opus model, covering both the Claude chat apps and the Claude Code terminal agent; the post went live on 22 September and reached Hacker News' front page. The short version: the model behaves differently enough that habits built around earlier Opus releases — small tasks, 'think step by step' nudges, constant check-ins — should be retired.
What changed in Opus 5.5
According to the guide, three shifts define the release. It sustains multi-part work on its own for much longer, with its largest gains over prior Opus models on multi-step tasks like carrying a change through a large repository until tests pass; early testers reportedly let it run coding jobs for hours with little oversight. It reports its work more plainly than Opus 5, stating what it did, what it found and what it still needs from the user. And it now reasons before every reply and calibrates the depth itself, which changes how people should prompt it.
Prompting: state the goal, drop the thinking nudges
The recommended pattern is to hand over the whole task in one message, define completion precisely — every endpoint migrated, the old client deleted, the test suite green — and name the only situations in which it should pause and ask. Users who remember something mid-run can type a follow-up and press Enter while Claude works, since restarting a long run now costs more.
Anthropic also advises deleting lines like 'think carefully' from prompts and saved instructions, because the model already reasons by default; in the company's own testing in a chat product, removing such a line made replies start sooner with no clear drop in quality. In Claude Code, reasoning depth should be tuned through the effort setting instead, and simple questions can be flagged for a direct answer.
For design work, the guide warns that the model falls back on a handful of default styles, and a vague instruction such as 'avoid a generic look' mostly swaps one default for another. It works better to list the specific patterns to exclude — the guide's example bars cream or off-white backgrounds, italic accent words, numbered section labels, monospace labels and pill-shaped buttons — and then append whatever the model picked anyway on the next attempt.
Steering long runs in Claude Code
Opus 5.5 posts status updates as it works, but it sometimes halts to report rather than continue: a summary that names its next step without taking it, or an offer to go on. Anthropic's remedy is a short CLAUDE.md rule telling it to keep going when a step needs no human input, fold status notes into its next action, and stop only when genuinely blocked or before anything destructive — deleting data, force-pushing, or changing anything outside the repository. The guide cautions that fewer stops mean users must keep their own guardrails, including permission prompts for destructive commands, and notes that pair programmers can invert the rule to request a one-line plan up front and a recap at the end; the model follows either instruction.
For audits and migrations spanning a large codebase, the guide suggests asking the model to fan the work out across parallel subagents, verify each one's evidence before accepting it, and finish with a single table of results. Because long runs fill the context window and Claude Code then summarizes older turns, it also recommends keeping the task checklist in a file that survives that compaction, updated as items close.
Checking the result
When a long run ends, the guide says to read what the model is waiting on — open decisions, changes needing approval — before the rest of the summary, and to fix the summary's shape via CLAUDE.md if the default format does not fit. One early tester cited by Anthropic said Opus 5.5 at its lowest effort setting caught more bugs than Opus 5 at high effort, with fewer false alarms, and that its plain-language explanations make pull request descriptions easier to review. For research tasks, asking the model to flag anything it could not confirm, and to say where it looked, surfaces the caveats worth reading. In the chat apps, Anthropic says to check the model picker and to attach charts, diagrams and screenshots rather than retyping their contents, since Opus 5.5 reads images more accurately than its predecessor and handles spatial relationships, such as which boxes an arrow connects.
Why it matters
The guide reads less like a benchmark announcement than a manual for a shift in where the human sits in the loop. When a model runs for hours and reasons unprompted, the leverage moves from per-message prompt engineering to task framing, explicit stop conditions and verification habits — file-based checklists, diff review and subagent evidence checks. Teams that treat Opus 5.5 like Opus 5 will over-prompt it, supervise the wrong things, or lose track of long runs; the guide is essentially a checklist for avoiding all three.
- #anthropic
- #claude-code
- #llm
- #prompting
- #developer-tools