claude-sonnet-4-6 — pricing, context window and limits
| Status | current |
|---|---|
| Input / MTok | $3 |
| Output / MTok | $15 |
| Context window | 1,000K tokens |
| Max output | 128K tokens |
| Min cacheable prefix | 1,024 tokens |
| Effort levels | low, medium, high, max |
| Thinking | adaptive-opt-in |
| Platforms | Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry |
The previous-generation Sonnet, still very much in active use
Sonnet 4.6 sits one generation behind the current Sonnet, at the same sticker price but with the older tokenizer — meaning identical text bills fewer tokens here than it would on Sonnet 5, despite the unchanged headline rate. That gap is exactly why a Sonnet 4.6 to Sonnet 5 migration needs to be re-baselined with a real token count rather than assumed cost-neutral.
Thinking here is opt-in, not default
Unlike Sonnet 5, this model only thinks when you explicitly ask it to — a real behavioural difference to account for if you're comparing the two generations side by side or building routing logic that treats them as interchangeable. It also tops out one effort tier below Sonnet 5's ceiling; the highest available tier doesn't exist on this generation at all.
Still a reasonable choice for stable, already-tuned workloads
For a workload that's already tuned and performing well on this generation, there's no urgency to move purely for its own sake — a deliberate migration, tested against real traffic, beats a rushed one every time, and this model remains fully active rather than approaching retirement.
Related
See migrating: Sonnet 4.6 to Sonnet 5 for the full field-by-field diff, and what actually changed from Sonnet 4.6 to Sonnet 5 for the narrative migration guide covering both the tokenizer and thinking-default changes in practical terms.
Verified 2026-08-08 against claude-api skill — shared/models.md + SKILL.md model table (canonical: https://platform.claude.com/docs/en/about-claude/models/overview, https://platform.claude.com/docs/en/pricing).