claude-opus-5-5 — pricing, context window and limits
The current Opus and Anthropic's default model, priced below Opus 5 with cache reads at 0.05x input. Thinking cannot be disabled — effort is the only control, and its API default is medium, one level below Opus 5's. Forced tool_choice returns a 400, and on the Claude API and Google Cloud computer use needs the computer toolset.
Price per million tokens
| Input | $4 |
|---|---|
| Output | $20 |
| Cache write, 5-minute TTL | $5 |
| Cache write, 1-hour TTL | $8 |
| Cache read | $0.20 (0.05x input) |
| Batch input / output | $2 / $10 |
| Fast mode input / output | $8 / $40 |
Limits and behaviour
| Status | Current |
|---|---|
| Context window | 1M tokens |
| Max output | 128K tokens |
| Min cacheable prefix | 512 tokens |
| Effort levels | low, medium, high, xhigh, max |
| Effort default | medium |
| Thinking | Always on — cannot be disabled |
| Forced tool_choice | Rejected (400) |
| Platforms | Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud, Microsoft Foundry |
| Released | 22 Sept 2026 |
Lifecycle
| API model name | claude-opus-5-5 |
|---|---|
| Retires no sooner than | 22 Sept 2027 |
The default model, priced below the Opus it replaced
Opus 5.5 is the current Opus and the model Anthropic's documentation treats as the default for most work, including complex agentic coding. It costs less per token than Opus 5, and its cache reads are discounted further than any model outside the Fable tier — the table above shows both. Anthropic's own framing is that the saving per solved task is larger than the per-token cut, because it tends to finish the same work in fewer tokens; that is a claim to verify on your traffic, not a number to budget against.
Thinking is always on, and effort is the dial
A disabled thinking configuration, or a fixed thinking budget, returns a 400 at every effort level.
Effort is now the only control over how much the model thinks — and its API default is medium,
one level below Opus 5's. A request that omits effort therefore runs lighter than it did on Opus 5,
while a request that keeps Opus 5's explicit setting tends to think more per turn than before,
especially at the top levels. Set effort explicitly and re-run the sweep rather than carrying a
value over. See Opus 5.5 defaults to medium effort.
Four breaking changes to check first
Disabled thinking is rejected. Forced tool choice (any or tool) is rejected. Thinking blocks
are bound to the model and conversation that produced them, so an edited history or a fallback to
another model loses them. And on the Claude API and Google Cloud, computer use works only through the
computer toolset — the older tool version returns a 400 there. None of these fail quietly, which is
the good news: each surfaces as an error in the first test run.
Refusals now include biology and reasoning extraction
Its safety classifiers cover more categories than Opus 5's. A decline arrives as a normal response with a refusal stop reason; ship a fallback from day one, and note that requests declined for trying to extract the model's reasoning are not retried on a fallback model.
Related
See migrating: Opus 5 to Opus 5.5, Opus 5.5 vs Sonnet 5.5 and what actually changed from Opus 5 to Opus 5.5.
Verified 2026-09-30 against claude-api skill — shared/models.md, model-migration.md + SKILL.md model table and Anthropic's pricing and model-deprecation pages (canonical: https://platform.claude.com/docs/en/about-claude/models/overview, https://platform.claude.com/docs/en/about-claude/pricing, https://platform.claude.com/docs/en/about-claude/model-deprecations).
Verified 2026-09-30 against Anthropic's model-deprecations status table and pricing page (canonical: https://platform.claude.com/docs/en/about-claude/model-deprecations, https://platform.claude.com/docs/en/about-claude/pricing).Could not confirm: These dates apply to the Claude API, Claude Platform on AWS and Microsoft Foundry. Amazon Bedrock and Google Cloud keep their own retirement schedules, which we do not track.