What actually changed from Opus 5 to Opus 5.5
Cheaper, and stricter about requests
Opus 5.5 keeps Opus 5's context window, output cap and tokenizer, so token counts don't move, and it costs less per token — the comparison shows both rate cards. Anthropic says existing Opus 5 prompts should perform well out of the box. The work is in the request code and in a handful of behaviours that change what a user sees. Here is what changed, starting with what fails.
Thinking can't be switched off
On Opus 5, thinking: {"type": "disabled"} was accepted at effort high or below. On Opus 5.5 there is no off switch: send a disabled config or a manual budget_tokens config and you get a 400, whatever the effort level. That leaves effort as the single dial for how much the model thinks, so a route that ran with thinking off should move to a low effort level instead — see
the error and its fix.
If a prompt asked the model to write out its reasoning in the reply as a substitute for thinking,
remove that instruction. On Opus 5.5 it can trigger a reasoning_extraction decline. Read the
thinking blocks with display: "summarized" instead.
The effort default dropped, and the levels moved
Leave effort unset and Opus 5.5 runs at medium, where Opus 5 ran at high. Yet level for level, Opus 5.5 usually thinks longer per turn than Opus 5, and the gap is widest at xhigh and max. So the
same setting doesn't mean the same behaviour in either direction. Set effort explicitly and re-run
your effort sweep rather than carrying the Opus 5 value across; the
effort default guide covers what each choice
does to the bill. To get less thinking, lower the level before adding "think less" instructions — it
works more reliably than prompting.
Forced tool choice is rejected
tool_choice of any or of a named tool returns a 400. Migrate by intent: leave tool_choice on
auto, use strict tool use to keep the schema guarantee, name the tool you expect in the prompt, and
check in code that the call happened. The error page has
the patterns.
Thinking blocks are bound to the conversation
Editing history between requests — a rebuilt system prompt, a trimmed tool result, a reminder inserted and later removed — now invalidates every later thinking block, and on newer accounts that is a 400. Opus 5.5 reads thinking blocks from Opus 5 and earlier models, so a conversation that moves onto it keeps its reasoning. The reverse doesn't hold: on the Claude API only Fable 5.1 and Mythos 5.1 can read an Opus 5.5 block, so a router switch or refusal fallback to Opus 5 continues without it. Making an agent loop safe for preserved thinking has the append-only replacements.
Computer use moves to the toolset
On the Claude API and Google Cloud, the older computer_20251124 tool returns a 400; Opus 5.5 accepts
only the computer toolset, which needs no beta header. Amazon Bedrock still accepts the older version.
Opus 5 accepts both, so make and test the change there first —
the error page shows both request shapes.
Progress notes arrive as thinking blocks
On Opus 5, the short notes the model writes between tool calls came back as text. On Opus 5.5, any
note longer than a sentence or two comes back as a progress-update thinking block — and under the
default display setting its text is empty. No request fails, but an interface that renders only text
goes silent for the length of a long agentic turn. Set thinking.display to "updates" (beta
thinking-display-updates-2026-08-18), render each non-empty thinking block before the tool call it
precedes, pass the blocks back unchanged, and say in the system prompt how often you want updates.
A broader set of refusals
Opus 5.5 adds a biology classifier to the cybersecurity one Opus 5 had, and it can decline requests that try to extract its internal reasoning. Declines arrive as an HTTP 200 with a refusal stop reason, and server-side fallbacks can retry them on another model — handling refusals walks through the setup.
Prompts worth revisiting
Opus 5-specific instructions. Prompts written to curb Opus 5's verbosity, over-verification or scope may no longer be needed. Keep them as a starting point, then test each one against your own evaluations instead of carrying it over untouched.
Frontend design. Asked for frontend work without design direction, Opus 5.5 falls back on a few default styles, and a general "avoid a generic AI look" mostly swaps one default for another. It responds well to a list of specific patterns to avoid; look at what the first result used and extend the list.
Visual inputs. It reads charts, diagrams and screenshots far more precisely out of the box, so cropping and zooming scaffolding built for earlier models may no longer earn its place. Re-test it before keeping it.
Long turns. At xhigh and max, turns run longer than on Opus 5. Plan timeouts, streaming and
progress indicators accordingly — and lower effort before prompting for brevity.
Verified 2026-09-30 against ClaudeHow facts module (src/data/facts/) — see /about/#accuracy.