Cache reads got cheaper on Fable 5.1 and Opus 5.5 — what changes
The discount stopped being one number
For as long as prompt caching has existed, every model discounted a cache read by the same fraction of input, so the only caching question was whether a prefix repeated often enough to earn back its write. The latest releases broke that symmetry. Fable 5.1 and Mythos 5.1 read their cache at a far deeper discount than any model before them, and Opus 5.5 sits between them and the rest of the lineup. Writes did not change at all. The all-models pricing table now shows the read rate beside every headline price for exactly this reason: two models with similar input rates can have very different read rates, and on a cache-heavy workload the read rate decides the bill.
What did not change: the break-even count
It is tempting to assume a cheaper read makes caching pay off sooner. It doesn't. For every read rate in the lineup, the break-even count is set by what a write costs, and writes kept their price. The counts on the prompt-caching reference apply to these models as printed — this site's test suite checks them against every model's actual read rate, precisely so that a pricing change like this one can't quietly invalidate them.
What does change is how much each hit is worth once you are past break-even, and how much each miss costs.
A miss now costs relatively more
On these models the gap between a cache read and a fresh input token is wider than it has ever been. Every request that misses — because the entry expired, because a timestamp crept into the system prompt, because a history edit changed the prefix — pays to write that prefix again at the cache-write rate, when it could have read it for a sliver of the price. The invalidators that used to be a nuisance are now expensive. Audit for them first: the silent invalidators table lists the patterns and why each one breaks the prefix.
Keep-alive can beat the long TTL
The usual advice for traffic whose gaps outlast the short TTL but not the long one was to pay for the
long TTL. For Fable 5.1 and Mythos 5.1, Anthropic's guidance is different: stay on the default short
TTL and, while the session is idle, re-send the previous request with max_tokens: 0 shortly before
the entry would expire. That request refreshes the entry's timer and bills only a cheap cache read,
with no output tokens. Because reads cost so little on these models, it usually works out cheaper
than the long TTL's pricier write — unless pauses regularly stretch toward the long TTL's full length.
The keep-alive has limits. Send it with streaming off: streaming is a transport option rather than
part of the cached prefix, so dropping it for that one request costs nothing. A request that uses
structured outputs can't be sent with max_tokens: 0, and neither can one inside a Message Batches
job. For traffic shaped like that, the long TTL is still the right tool.
Model comparisons now need the read column
A comparison built on input and output rates alone now misleads in a specific way. Opus 5.5 and Sonnet 5.5 have identical cache-read rates, even though Opus 5.5's input rate is higher, so for an agent that mostly re-reads a cached history, the input side of the gap between them largely disappears. Output still costs what the rate card says. Fable 5.1's read rate sits close to both. Price your own token mix in the token & cost estimator and the prompt-caching savings calculator, both of which use each model's own read rate.
Check that the reads are happening
None of this matters if the cache isn't being read. After any change to how prompts are assembled,
check the usage block: reads show as cache_read_input_tokens, writes as
cache_creation_input_tokens. On a healthy multi-turn loop, reads grow turn over turn and writes stay
small. If writes are close to the full conversation size on every turn, something upstream of the
breakpoint is changing — and on these models, that costs more than it used to.
Verified 2026-09-30 against ClaudeHow facts module (src/data/facts/) — see /about/#accuracy.