Sonnet 5.5 vs Haiku 4.5
Side by side
| Sonnet 5.5 | Haiku 4.5 | |
|---|---|---|
| Input / MTok | $2 | $1 |
| Output / MTok | $10 | $5 |
| Cache read / MTok | $0.20 | $0.10 |
| Context window | 1M | 200K |
| Max output | 128K | 64K |
| Min cacheable prefix | 512 | 4,096 |
| Effort levels | low, medium, high, xhigh, max | none |
| Effort default | high | — |
| Thinking | On by default — lowest setting is between_tools | Fixed budget_tokens, off unless requested |
| Forced tool_choice | 400 | accepted |
The choice for high-volume work
This pairing decides most high-volume workloads, where per-token cost compounds across enormous request counts. Haiku 4.5 is the cheaper per-token option; Sonnet 5.5 costs more but brings real reasoning depth, the full effort range and a much larger context window.
The cache minimum can flip the arithmetic
The comparison most people run stops at the headline rate. The row that deserves more attention is the minimum cacheable prefix. Haiku 4.5 has one of the highest in the lineup and Sonnet 5.5 one of the lowest, so a workload built on a short, stable system prompt can cache on Sonnet 5.5 and silently fail to cache on Haiku 4.5. When most of each request is that prefix, the cheaper model can end up billing the prefix in full while the dearer one bills it as cheap cache reads. Price your real prompt size on both before assuming Haiku wins.
Request shape differs, not just price
Haiku 4.5 rejects the effort parameter and uses the older fixed-budget thinking; Sonnet 5.5 thinks
by default and accepts every effort level. A shared request builder has to branch between them. Haiku
also carries the smaller context window and output cap shown in the table, which rules it out for
anything that has to hold a large codebase or document set in one request.
Plan for the lifecycle, too
Haiku 4.5's earliest possible retirement date is the closest of any current model's. That isn't a retirement, but a pipeline pinned to it should know where it would go — and routing the easy majority to Haiku with the harder minority on Sonnet 5.5 keeps that move small.
Related
See claude-haiku-4-5, the minimum-cacheable-prefix table and Claude model retirement dates.
Verified 2026-09-30 against claude-api skill — shared/models.md, model-migration.md + SKILL.md model table and Anthropic's pricing and model-deprecation pages (canonical: https://platform.claude.com/docs/en/about-claude/models/overview, https://platform.claude.com/docs/en/about-claude/pricing, https://platform.claude.com/docs/en/about-claude/model-deprecations).Could not confirm: On Microsoft Foundry it is hosted on Azure only, so Foundry features that require Anthropic hosting (code execution, the Files API, the newer web tools) are unavailable for it there.