Opus 5.5 defaults to medium effort: what that does to your bill
One field you probably never set
Most code never sets output_config.effort. On Opus 5 and the Opus models before it, leaving it out
meant high. On Opus 5.5 it means medium — and Claude Code runs both Opus 5.5 and Sonnet 5.5 at
medium by default too. Nothing errors when a default changes. The model simply thinks at a different
level than it did, and your bill moves without any line of your code moving.
What happens to code that omits effort
For code that moves from Opus 5 without touching effort, the direction is usually cheaper. Anthropic
reports that on many coding, analysis and vision tasks, Opus 5.5 at its default effort matched or beat
Opus 5 while using fewer tokens — and it costs less per token too. Fewer tokens matters more than it
sounds, because thinking bills as output even when its text isn't returned to you. In Anthropic's
testing, Opus 5.5 at medium exceeded Opus 5 at high on coding and knowledge-work evaluations.
That describes their evaluation set, not your prompts. The honest summary: expect a lower bill, and
check that the outputs you care about didn't quietly get shallower.
What happens to code that sets effort explicitly
Code that pinned high on Opus 5 and carries that value across gets the opposite surprise. At a given
level, Opus 5.5 tends to think more per turn than Opus 5 did, especially at xhigh and max. So an
explicit high doesn't mean "the same amount of thinking as before" — it can mean longer turns and
more output tokens than your Opus 5 baseline. Effort names don't describe the same amount of thinking
from one model to the next, which is why copying a setting across by name is the wrong comparison.
The right comparison: cost per solved task
Effort trades tokens per request against the chance that a request finishes the job. A cheaper request that needs a retry, a follow-up turn or a human fix isn't cheaper. Anthropic frames Opus 5.5's economics as cost per solved task, and that is the right frame for your own measurement: run a sample of real tasks at neighbouring effort levels and compare the total spend per task that met your bar, not the spend per request. The effort & thinking cost estimator gives the per-request curve; your own evaluation supplies the success rate that turns it into a cost per task.
Set it explicitly, even if you keep the default
Anthropic's migration guidance is to set effort explicitly on Opus 5.5 and re-test, rather than
inheriting whatever the default happens to be. An explicit value is one you chose, and it protects the
route from the next default change. Choose per route rather than globally: short, scoped work such as
classification, extraction and chat usually holds up at low, while intelligence-sensitive work is
where the higher levels earn their cost. If you want less thinking, lower the effort level before you
add "think less" instructions — it reduces thinking, cost and latency more reliably than prompting.
Varying effort without paying for it twice
If one conversation needs different depths at different moments, don't change the top-level effort between requests: that invalidates the messages cache, and on Opus 5.5, with its deep cache-read discount, a cache miss is an expensive event. Opus 5.5 supports per-message effort instead — a system message inside the conversation that changes effort from that point on without invalidating the cache. See per-turn effort for the request shape and the models that accept it.
Size max_tokens for thinking you can't turn off
One practical trap comes with the change. Opus 5.5's thinking can't be disabled, and it counts toward
max_tokens. A limit sized for Opus 5 with thinking off can cut replies off after the reasoning.
Raise it before you compare quality, or you'll be measuring truncation rather than effort.
Verified 2026-09-30 against ClaudeHow facts module (src/data/facts/) — see /about/#accuracy.