ClaudeHowSupport Us

What actually changed from Sonnet 5 to Sonnet 5.5

Same rate card, different request rules

Sonnet 5.5 keeps Sonnet 5's per-token prices and its tokenizer, so identical traffic produces an identical invoice — or a smaller one, because the minimum cacheable prompt dropped and shorter prompts now cache (the cache minimums table has both values). Anthropic says existing Sonnet 5 prompts should perform well without changes. What doesn't carry over cleanly is the request code: several things that worked on Sonnet 5 now return a 400.

Disabled thinking is gone; between_tools is the floor

On Sonnet 5, thinking: {"type": "disabled"} turned thinking off. On Sonnet 5.5 it returns a 400 whose message points you at the replacement. thinking: {"type": "between_tools"} is now the lowest setting: the model does no extended thinking, and the short progress notes it writes between tool calls come back as thinking blocks with their text included. It comes with limits, each one a 400 when crossed:

  • Effort must be high or below. At xhigh or max, use adaptive thinking.
  • Nothing else can go inside thinking — no display, budget_tokens or block_binding.
  • Effort can't change partway through the conversation.
  • No other model accepts it, which means a router or retry that forwards the same request body elsewhere has to strip it out first.

Anthropic's recommended migration order is to try adaptive thinking at low effort first — at low the model keeps thinking short and skips it on most simple requests — and to measure latency and quality on your own traffic. Use between_tools only if that doesn't hold up. Either way, delete any instruction telling the model not to think; those make it more likely to leak internal XML tags into visible output. The error page has the exact messages.

The other request changes

Forced tool choice. tool_choice of any or a named tool returns a 400, including on the token-counting endpoint. See the fix.

Preserved thinking. Thinking blocks are bound to the conversation that produced them, and editing earlier history invalidates every later block. New accounts are enforced by default on the Claude API and Amazon Bedrock. The binding controls work only with adaptive thinking, so a between_tools route has to stay append-only, or strip its thinking blocks from any edited turn onward. See making an agent loop safe for preserved thinking.

Computer use. On the Claude API and Google Cloud, the older computer tool version returns a 400; use the computer toolset. See the error page.

The advisor tool. A Sonnet 5.5 executor rejects Opus 4.8, Opus 4.7, Opus 4.6, Sonnet 5 and Sonnet 4.6 as advisors. The accepted advisors return their advice encrypted, so its text isn't readable in the response.

Effort levels were recalibrated

high remains the API default, yet the same level name now buys a different amount of thinking than it did on Sonnet 5 — so repeat your effort sweep and pin the level explicitly. Where to start, per Anthropic: medium for agentic coding and multi-step tool use; low for chat, content generation, classification, extraction and search. From medium up, the model thinks briefly before almost every reply — even a greeting — and a system-prompt request to think less has almost no effect; lowering the level is the lever that works.

Judge the change by cost per completed task, not per token. In Anthropic's testing, Sonnet 5.5 finished agentic coding and multistep tool-use work in far fewer requests than Sonnet 5, and on most agentic coding evaluations it scored higher at medium than Sonnet 5 did at high, at a fraction of the cost. To vary effort within a conversation without resetting the cache, use per-message effort — it needs adaptive thinking.

Progress notes and new capabilities

As on Opus 5.5, notes longer than a sentence or two between tool calls now arrive as progress-update thinking blocks, empty under the default display. With adaptive thinking, set thinking.display to "updates" to get them as text, and render them before the tool call that follows.

Several features Sonnet 5 never had arrive with it: mid-conversation system messages, mid-conversation tool changes, per-message effort, task budgets, and tool definitions inside a mid-conversation message. It also declines in more categories than Sonnet 5 — see handling refusals for which ones a fallback retries.

Prompt changes worth making

Remove workarounds for what improved. It uses connected tools more reliably, declines fewer benign requests and holds a system-prompt role better. Refusal steering, tool-call retry shims and "do not be lazy" instructions written for Sonnet 5 can go — remove them and re-run your evaluations before tuning anything else.

Stop discouraging tools. It follows "only use tools when strictly necessary" literally, and on chat and knowledge work it can hold off on tools until asked. Delete that kind of language where your product should prefer connected sources.

Deliver mid-task user input as a user turn. A user's message placed directly after a tool result, or inside one, can be read as a prompt-injection attempt. Put it in a text block after the last tool result, and keep harness notices in a separate system message.

Ask for real verification at low effort. Mostly at low, it can report a code change as done without running a check that exercises it. If a coding agent runs at low, tell it to run the project's tests, type-checker or build before reporting a change complete.

Be tolerant of near-miss tool names. It occasionally calls a tool by a name that differs only in letter case. Accept the call when the match is unambiguous, or return an error result stating the exact expected name — it usually corrects itself on the next turn.

Verified 2026-09-30 against ClaudeHow facts module (src/data/facts/) — see /about/#accuracy.