Migrating: Fable 5 to Fable 5.1
What changes
| From | To | |
|---|---|---|
| Model | Fable 5 | Fable 5.1 |
| Input / MTok | $10 | $10 |
| Output / MTok | $50 | $50 |
| Cache read / MTok | $1 | $0.25 |
| Min cacheable prefix | 512 | 512 |
| Context window | 1M | 1M |
| Effort levels | low, medium, high, xhigh, max | low, medium, high, xhigh, max |
| Effort default | high | high |
| Thinking | Always on — cannot be disabled | Always on — cannot be disabled |
| Forced tool_choice | accepted | 400 |
The same price, a different cache bill
Fable 5.1 keeps Fable 5's per-token price for input and output, its tokenizer, context window, output cap and cache minimum. The line that changes is cache reads, which now cost a fraction of what they did — the table above shows both rates. For an agent that re-reads a large prefix every turn, cache reads can be most of the input bill, so this is a real saving — Anthropic estimated typical workloads would cost noticeably less wherever usage is billed by the token. It also shifts the economics of keeping a cache warm: a miss now costs far more relative to a hit, so a harness that lets its cache lapse between turns is leaving more money on the table than it did.
The breaking changes
Forcing a tool with any or a named tool now fails with a 400 — on the Messages API, in batches and on the token-counting endpoint alike. A Fable 5.1 thinking block is readable by Mythos 5.1 and by nothing else, so when a request falls back to another model, that model starts without the reasoning. And editing earlier turns invalidates every
later thinking block — a 400 for accounts created on or after the enforcement date, on every
platform. Claude Code, claude.ai and the Agent SDK already keep histories append-only; code that
builds its own message array should be checked before it moves.
Two operational differences
Fable 5.1 has no Priority Tier, which Fable 5 does — a caller relying on it loses it on migration. And the two models share one rate-limit pool, so running both side by side during a migration spends the same headroom twice. Plan the cutover with that in mind rather than doubling traffic.
Related
See making an agent loop safe for preserved thinking and cache reads got cheaper on Fable 5.1 and Opus 5.5.
Verified 2026-09-30 against claude-api skill — shared/models.md, model-migration.md + SKILL.md model table and Anthropic's pricing and model-deprecation pages (canonical: https://platform.claude.com/docs/en/about-claude/models/overview, https://platform.claude.com/docs/en/about-claude/pricing, https://platform.claude.com/docs/en/about-claude/model-deprecations).Could not confirm: Availability on Amazon Bedrock, Google Cloud and Microsoft Foundry is not stated in the source we used, so we do not claim it either way.Could not confirm: Microsoft Foundry availability is stated for Anthropic-hosted deployments only, and the 1M context window on Amazon Bedrock and Google Cloud was still open at launch, so we do not promise either.