ClaudeHowSupport Us

Opus 5.5 is cheaper per token. That is not the interesting part.

The headline everyone will quote

Opus 5.5 costs less per token than Opus 5, for input and output alike, and the comparison page shows both rate cards side by side. The natural move is to take last month's Opus 5 bill, scale it down by the ratio, and call that the saving. Anthropic's own migration guidance warns against exactly that: re-baseline your costs rather than scaling them by the list price alone. It's good advice, because the per-token cut is the least interesting of the four changes that decide what Opus 5.5 actually costs you.

First: the cache read fell much further

The cut to cache reads is far deeper than the cut to input. Opus 5.5 reads its cache at a smaller fraction of input than Opus 5 did — and than every model outside the Fable 5.1 family still does. For an agent that re-reads a long cached history on every turn, cache reads are most of the input bill, so for that workload the effective saving is much larger than the headline suggests.

There are two side effects. Next to a hit, a miss is now far more expensive, so the silent invalidators that used to be a rounding error — a timestamp in the system prompt, a tool list rebuilt per request — now show up on the invoice. And Opus 5.5 reads cache at exactly the same rate as Sonnet 5.5. On a workload dominated by cached input, most of the input-side premium for choosing Opus disappears, and the real difference between the two comes down to output. That is a genuinely new shape for the Opus-versus-Sonnet decision, and nothing on a two-column rate card reveals it.

Second: the default effort dropped

Leave effort unset and Opus 5.5 runs at medium; Opus 5 ran at high. Anthropic reports that at its default effort, Opus 5.5 matched or beat Opus 5 on many coding, analysis and vision tasks while using fewer tokens — and on multistep coding work in a real repository, it matched or beat Opus 5's high-effort results in fewer steps. Fewer tokens at a lower price per token compounds. For code that never set effort, this default change may save more than the price cut does.

It cuts the other way for code that pinned an effort level. Hold the level constant and Opus 5.5 generally does more thinking per turn than its predecessor, most noticeably at the top of the range. Carry high across by name and you can get longer turns and more output tokens than your baseline. The effort guide has the details; the short version is to set effort explicitly and re-measure.

Third: thinking can't be turned off

On Opus 5, a route could disable thinking at high effort or below. On Opus 5.5, thinking is always on, and a disabled config is a 400. For a route that ran Opus 5 without thinking — a classifier, an extraction step, a latency-sensitive endpoint — this is the one place the bill can go up rather than down, because that route now pays for thinking tokens it never paid for before. Low effort keeps thinking short, so that is where those routes should start. Size max_tokens with room for the thinking too; a limit tuned for thinking-off can cut answers off.

Fourth: the right unit is a solved task

Anthropic's own yardstick for Opus 5.5 is what a solved task costs, and it is the right yardstick. A cheaper request that has to be retried, followed up or corrected by a person isn't cheaper. On knowledge work, Anthropic describes Opus 5.5 as much less likely than Opus 5 to state a figure or cite a source its inputs don't support, and a wrong figure that someone has to catch costs more than any token. On visual inputs it reads charts and screenshots far more accurately, so scaffolding built to crop and re-check images on earlier models may simply be unnecessary. Those savings never show up in a per-token comparison, which is precisely why per-token comparisons undersell this release.

What to expect on your own bill

Code that omits effort and leaves thinking on. Probably down, and by more than the headline ratio. Verify quality on your own evaluations rather than trusting that it held.

Cache-heavy agents. Down the most, provided the cache actually hits. Check cache_read_input_tokens after the switch and hunt down any invalidator before comparing costs.

Routes that ran with thinking disabled. Possibly up. Test them at low effort and compare cost per solved task, not per request.

Routes that pinned a high effort level. Unclear until measured — the level means more thinking than it did on Opus 5.

Whichever group you're in, re-baseline instead of scaling. The token & cost estimator and the effort & thinking cost estimator give the per-request side; your own evaluation set supplies the success rate that turns it into a cost per task.

Verified 2026-09-30 against ClaudeHow facts module (src/data/facts/) — see /about/#accuracy.