Opus vs Sonnet: what's the real gap in your Claude Code bill?
See the sticker-price gap between the two models, then the honest number for what pinning just the Sonnet-safe turns actually saves - next to the naive number most routing pitches quote instead.
A clean setup is just the beginning.
UsageCut trims that for free. Want Claude Code to be genuinely better, not just leaner? ClockedCode is the curated, conflict-free setup - installed in one paste.
Upgrade your Claude Code- 10x your Claude Code with hand-picked tools
- Specialist agents on your team
- Stop spending hours researching plugins
Two numbers, and why the smaller one is the honest one
Every calculator that compares Opus and Sonnet answers the easy question: multiply your tokens by each model's rate and subtract. That part is simple - Sonnet 5 currently runs a uniform 40%of Opus 4.8's rate on every axis (input, cache write, cache read, output), the same rate card behind the full Claude Code cost-per-token breakdown. This tool shows that gap first, for your current mix of Opus and Sonnet against running everything on one model.
The harder, more useful question is what actually pinning some of your work to Sonnet saves - and that number is smaller than the easy math suggests. This project's own model-routing research found roughly 39.8% of turns are "Sonnet-safe," but those turns carry only about 14.4% of a session's output tokens, and switching models mid-thread triggers a cache rebuild that erases the saving if done naively. So this tool shows a naive estimate (39.8% of your Opus spend, struck through) next to an honest one, credited only against the output-token share those turns actually carry - the same cache-safe reasoning behind the subagent cost calculator, which sizes an actual pinned subagent call by call.
This is a modeled comparison from documented per-token rates and this project's own measured session composition, not a reading of your actual transcripts. Run the free scan below for your real, per-session Opus/Sonnet split.
More free tools for Claude Code
Small, focused, no-signup tools from the maker of ClockedCode. Free to use, forever.
More free Claude Code tools are on the way - built one at a time by a solo dev who ships daily.
Claude Code tips, every Sunday
One short email a week: the token-saving tricks, setup tweaks, and tools worth your time. Free, unsubscribe anytime.
About this estimate
Is Sonnet actually 40% of Opus's price, or is that rounded?
It's the real rate card, not a rounding. Under Sonnet 5's current introductory pricing (through 2026-08-31), Sonnet lands at exactly 40% of Opus 4.8 on every axis - input, 5-minute cache write, cache read, and output - so a token costs 60% less the moment it runs on Sonnet instead of Opus. See the full input/output/cache-write/cache-read rate card in the Claude Code cost-per-token breakdown.
Why isn't the subagent-pinning savings close to 39.8%?
Because 39.8% is a share of turns, not a share of cost, and this project's own model-routing research found those Sonnet-safe turns carry only about 14.4% of a session's output tokens - easy turns are cheap turns. Worse, modeling per-turn routing honestly (with the real cache-miss penalty a model switch triggers) actually loses money: -122%. The naive number this tool shows next to the honest one is there specifically to make that gap visible, not to sell you on it.
So does pinning to Sonnet ever pay off?
Yes, but only through cache-safe coarse routing - a whole subagent pinned to Sonnet or Haiku with its own separate cache, not a switch mid-thread. That's the same mechanism the subagent cost calculator prices. The realistic envelope for that lever, measured across this project's own research, runs low single digits up to roughly 10-15% of total spend, materially larger only if you're currently defaulting everything to Opus.
Where does the 'share of tokens on Opus today' number come from if I don't know mine?
100% is a reasonable default, not a scare tactic - Claude Code custom subagents default to Opus unless a frontmatter model field or CLAUDE_CODE_SUBAGENT_MODEL override says otherwise, and most setups never add one. If you already run some work on Sonnet or Haiku, lower the slider; the honest-savings math scales down with it, since there's less Opus spend left to shave.
Is this my real Opus vs Sonnet bill?
No - it's a modeled comparison built from Anthropic's published per-token rates and this project's own measured session composition (1,037 real sessions), not a reading of your actual transcripts. Your real split depends on which subagents and models you actually invoke. Run the free UsageCut scan for your own numbers instead of an estimate.
Questions, answered
Is UsageCut really free?
Yes. The scan and every fix are free - no card, no subscription, no signup. You only give an email if you want the optimization plan and one-command undo sent to your inbox.
Is my code safe? Does UsageCut upload anything?
The scan runs entirely on your machine. It never uploads your code, never reads your API key, and never sends your prompts or conversations anywhere. Only anonymous summary counts leave, and only if you choose to share them.
How does it reduce Claude Code token usage?
It reads your local setup and session history to find waste - idle MCP servers loaded into every session, an oversized CLAUDE.md re-sent on every request, duplicate file reads, and bloated tool output - then trims it losslessly, so each session carries less context and you hit your limit far less often.
What is ClockedCode?
ClockedCode is the curated, conflict-free Claude Code setup from the same maker. UsageCut makes your current setup leaner; ClockedCode makes it genuinely better - a set of hand-picked tools, specialist agents, and a tuned CLAUDE.md, installed in one paste.
Who is ClockedCode for?
Developers who use Claude Code daily and want a setup that is powerful out of the box instead of spending hours researching and wiring up plugins, agents, and MCP servers themselves.
What's included in ClockedCode?
A vetted, conflict-free bundle of hand-picked tools, specialist agents, and a tuned CLAUDE.md - everything installed together in one paste, with no setup archaeology.
Is ClockedCode a one-time payment?
Yes. ClockedCode is a one-time purchase with an instant download - no subscription and no recurring fees.
Does ClockedCode work with my existing setup?
Yes. It installs alongside what you already have and is fully reversible. Run UsageCut first to clean your current setup, then add ClockedCode to level it up.