Claude Sonnet 5 Costs Less Than Opus—Until It Doesn’t

Pick the wrong effort level on Sonnet 5 and you will pay more than Opus 4.8 for nearly identical output.

Mid-tier model pricing stopped being predictable

Running long agentic coding tasks through a model that silently escalates token spend is a budget problem disguised as a capability upgrade. Engineers and AI leads need a clear map of where Sonnet 5 wins on cost and where it quietly crosses the Opus 4.8 line.

Sonnet 5 is a reasoning engine with a throttle, not a fixed-rate API call

Anthropic Claude Sonnet 5 vs Sonnet 4.6 vs Opus 4.8 ships with four effort levels: low, medium, high, and xhigh, each spending progressively more tokens on internal reasoning before returning output. You call the API exactly as before, but you now control a multiplier that changes both quality and cost per request. The catch is a new tokenizer, shared with Opus 4.7, that maps the same input text to roughly 1.0 to 1.35 times more tokens than before.

Coding teams running autonomous agents feel this first

  • Senior engineers managing Claude Code pipelines who need SWE-bench Pro accuracy at 63.2% without Opus 4.8 spend
  • AI infrastructure leads benchmarking browser and terminal automation, where Claude Sonnet 5 hits 81.2% on OSWorld-Verified
  • Platform architects on Team or Enterprise plans who set default models across dozens of users and absorb every tokenizer-driven cost shift

At low and medium effort, Claude Sonnet 5 is the obvious call. At xhigh, the math changes fast.

Anthropic moved the mid-tier ceiling right when the market expected a flagship drop

Intro pricing runs $3 per million input tokens and $15 per million output tokens after August 31, compared to $5 and $25 for Opus 4.8. That gap closes quickly at xhigh effort on long agentic sessions, meaning the real competitive pressure lands on GPT-4o and Gemini 1.5 Pro at the mid tier, not on Opus.

What you can actually do with it today

  • Run multi-step browser automation tasks with better self-correction on tool failures
  • Deploy inside Claude Code for extended coding sessions without context loss
  • Benchmark against Opus 4.8 on HLE tasks at 57.4% accuracy before committing to flagship spend
  • Set effort level per request to cap token spend on low-stakes subtasks

Intro pricing of $2/$10 per million tokens holds through August 31, 2025, then rises to $3/$15.

The tokenizer change is the detail most teams will miss until the invoice arrives

The updated tokenizer means existing cost estimates built on Sonnet 4.6 data are wrong by up to 35%. That is a real limitation that requires re-benchmarking any production pipeline before treating intro pricing as a long-term budget figure.

GPT-4o remains the default mid-tier choice for teams already inside the OpenAI ecosystem. Gemini 1.5 Pro undercuts on price but trails on agentic reliability benchmarks where Claude Sonnet 5 now leads.

Mid-tier AI pricing just became a variable, not a line item

The effort-level architecture means model cost is now a runtime decision, not a procurement one — and most finance and engineering teams are not set up to govern that yet. We cover tools like this every Friday — subscribe here and we’ll send the best ones straight to you.