ClaudeHowSupport Us

claude-sonnet-5-5 — pricing, context window and limits

The current Sonnet, at Sonnet 5's prices and tokenizer, with the prompt-cache minimum halved to 512 tokens. thinking: disabled now returns a 400 — the lowest setting is between_tools, accepted only at effort high or below — and forced tool_choice returns a 400. Effort still defaults to high, but the levels are recalibrated, so re-run your effort sweep.

Price per million tokens

Input$2
Output$10
Cache write, 5-minute TTL$2.50
Cache write, 1-hour TTL$4
Cache read$0.20 (0.1x input)
Batch input / output$1 / $5

Limits and behaviour

StatusCurrent
Context window1M tokens
Max output128K tokens
Min cacheable prefix512 tokens
Effort levelslow, medium, high, xhigh, max
Effort defaulthigh
ThinkingOn by default — lowest setting is between_tools
Forced tool_choiceRejected (400)
PlatformsClaude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud, Microsoft Foundry
Released28 Sept 2026

Lifecycle

API model nameclaude-sonnet-5-5
Retires no sooner than28 Sept 2027

Same price as Sonnet 5, and it caches shorter prompts

Sonnet 5.5 is the current Sonnet. It kept Sonnet 5's per-token prices and tokenizer, so a move between them costs nothing on the rate card and needs no token recount. What it changed on the cost side is caching: its minimum cacheable prefix is lower than Sonnet 5's, so system prompts that were too short to cache before can start producing cache reads with no code change. Check the usage block after switching — it is one of the few migrations where the bill can fall on its own.

The off switch moved

thinking: {type: "disabled"} returns a 400 here. The lowest setting is now between_tools, which does no extended thinking and returns the model's short progress notes between tool calls. It comes with limits of its own: it is accepted only at effort high or below, takes no other field inside thinking, cannot be combined with a per-message effort change, and exists on this model only — any other model rejects it. Anthropic's advice is to try adaptive thinking at low effort before reaching for it. See disabled thinking rejected on Sonnet 5.5.

Effort levels were recalibrated

The default is still high, but a level no longer produces the same amount of thinking it did on Sonnet 5. The documented starting points are medium for agentic coding and multistep tool use, and low for chat, extraction and classification. Re-run your effort sweep instead of reusing the Sonnet 5 setting — the level names survived the change, the meaning didn't.

The rest of the breaking changes

Forced tool choice returns a 400. Thinking blocks are tied to this model and conversation — no other model reads them. Computer use on the Claude API and Google Cloud needs the computer toolset. And the advisor tool accepts fewer advisor models. Declines arrive in more categories than on Sonnet 5, and server-side fallback retries only some of them.

See migrating: Sonnet 5 to Sonnet 5.5 and Sonnet 5.5 vs Haiku 4.5.

Verified 2026-09-30 against claude-api skill — shared/models.md, model-migration.md + SKILL.md model table and Anthropic's pricing and model-deprecation pages (canonical: https://platform.claude.com/docs/en/about-claude/models/overview, https://platform.claude.com/docs/en/about-claude/pricing, https://platform.claude.com/docs/en/about-claude/model-deprecations).Could not confirm: On Microsoft Foundry it is hosted on Azure only, so Foundry features that require Anthropic hosting (code execution, the Files API, the newer web tools) are unavailable for it there.

Verified 2026-09-30 against Anthropic's model-deprecations status table and pricing page (canonical: https://platform.claude.com/docs/en/about-claude/model-deprecations, https://platform.claude.com/docs/en/about-claude/pricing).Could not confirm: These dates apply to the Claude API, Claude Platform on AWS and Microsoft Foundry. Amazon Bedrock and Google Cloud keep their own retirement schedules, which we do not track.