Claude Sonnet 5.5

Same-Task Receipts Against Opus 5.5, and When the Cheaper Model Is Enough

EffortRoutingAcceptance checkRefusalCost per accepted output

Six days after Opus 5.5, Anthropic shipped the cheaper half of the pair. Sonnet 5.5 costs what cost, which is half of per token, and the launch post calls it faster and cheaper per task than its predecessor. The useful question is not whether it is good. It is which of your jobs it can take from Opus without you noticing, and which ones you would notice. So I gave both models, and Sonnet 5, the same job on the same day and counted what they found.

Identify the model before judging it#

Sonnet 5.5 was released on 28 September 2026 as the second model of the 5.5 family. The API ID is claude-sonnet-5-5 (on Amazon Bedrock, anthropic.claude-sonnet-5-5), and like every Claude ID since the 4.6 generation it is dateless but pinned to one snapshot. It has a 1M-token context window, a 128K maximum output (300K on the Batches API with the output-300k-2026-03-24 beta header) and a June 2026 knowledge cutoff. Anthropic’s lineup table labels its latency “Fast” against Opus 5.5’s “Moderate”, a relative label rather than a measurement. Sonnet 5 is not deprecated. Model overview, models overview.

What Anthropic claims, and the conditions#

The price is Sonnet 5’s: $2 per million input tokens, $10 output, $0.20 for cache reads. Cache writes are $2.50 for five minutes and $4 for an hour, and the Batches API is $1 and $5. The launch post says it “generates outputs 30%+ faster than Sonnet 5” and, “in our testing, it costs up to 30% less per task”. Neither claim comes with a workload, an effort level or a method. Anthropic’s product page calls the cost figure “an estimated” one. Launch post, pricing.

The defaults differ by surface, and that changes every comparison. “In Claude Code and our apps, the default effort is set to Medium, while the Claude Platform defaults to High.” Opus 5.5’s API default is medium, so two API calls that both omit effort run Sonnet 5.5 one level higher than Opus 5.5.

The system card gives the standard configuration behind these rows: max effort, averaged over five trials, unless noted.

BenchmarkSonnet 5.5Sonnet 5Opus 5.5Conditions
Terminal-Bench 4.070.6%10.3%66.4%Opus 5.5 at xhigh (64.8% at max). Sonnet 5’s score appears only in the launch post, unexplained
FrontierCode v1.146.2%42.4%54.4%Run by Cognition. Sonnet 5.5 scores 52.1% at xhigh; at max it more often ran the code-review skill’s subagents, which in two cases Cognition examined led to a timeout or edits beyond the task
CursorBench 4.055.5%34.1%57.8%Run by Cursor
GDPval-AA v2.1184414491846Elo, run by Artificial Analysis on a pre-release Sonnet 5.5 deployment with a since-fixed structured-output bug
OSWorld 2.180.1%57.0%81.8%Partial credit. Strict pass rate: 43.5%, 25.6%, 48.7%
SWE-bench Pro81.3%63.2%89.9%System card only

The Terminal-Bench row is the one to treat with care. Sonnet 5 scores between 3.2% and 10.3% at every effort level in the launch chart, and neither the post nor the system card explains why. My measurement below shows a large efficiency change from Sonnet 5 to 5.5. It does not explain a jump of that size.

The more useful data is behind the charts. For four benchmarks the launch page plots score against cost per task at each effort level, and the numbers are in the page. Sonnet 5.5’s costs on FrontierCode and CursorBench are Anthropic’s own estimates from token counts at list prices: the system card says Cognition reported no cost for it, and Cursor’s published costs cover the other models. At each model’s API default, Sonnet 5.5 at high and Opus 5.5 at medium, Opus 5.5 scores higher on all four and costs 1.1 to 1.9 times as much per task:

BenchmarkSonnet 5.5 at highOpus 5.5 at mediumSonnet 5.5 at xhigh
Terminal-Bench 4.0, per attempt43.0% at $1.9457.6% at $2.9461.5% at $5.30
FrontierCode v1.149.4% at $0.4254.6% at $0.8052.1% at $1.59
CursorBench 4.047.8% at $1.6752.5% at $2.9153.1% at $3.88
AA-Briefcase v1.11634 at $3.951642 at $4.401746 at $9.63

Read across a row and the trap appears. Pushing Sonnet 5.5 up to xhigh to catch Opus often costs more than Opus at medium: on FrontierCode, Opus at medium scored higher for half the money. At max, Sonnet 5.5 can cost more per task than Opus 5.5 at max: $29.19 against $21.05 on AA-Briefcase, and $20.78 against $6.19 on FrontierCode. The launch post says as much: “At higher settings, it can perform comparably at a similar cost.” Cheap per token is not cheap per task at every setting.

The other headline claims have no stated conditions. Sonnet 5.5 is “the first Sonnet model to beat Pokémon Red working only from screenshots”, with no harness, step count or effort level given. Its slide-deck example is one internal test judged by two experts. The 13 customer quotes are vendor-published anecdotes.

What independent boards found on day one#

Captured on 29 September, the day after launch. Most boards had not listed Sonnet 5.5 yet: not LMArena, ARC Prize, SimpleBench, the official Terminal-Bench board or Epoch’s capability index. The ones that had agree on one shape.

What I measured: one review, three models#

On 29 September I gave Sonnet 5.5, Opus 5.5 and Sonnet 5 the same job. The file was a 209-line Python script that retires git worktrees, with six bugs planted in it, and the task was to report every place where the code breaks its own docstring or could lose data. Every run got the same prompt and file, as a headless Claude Code session (claude -p --model <id> --effort <level>) allowed only to read files and run python3 -c, two runs per setting. Code graded each report against the answer key: a finding within two lines of a planted bug counts as found. The fixture comes from the effort sweep I ran on 22 September (Chapter 54), so it existed before Sonnet 5.5 did. Claude Code recorded the model that served every message, and it was always the one requested.

Same review, three models: output tokens, time and bugs found
Same review, three models: output tokens, time and bugs found Six planted bugs, same prompt, two runs per setting, 29 Sep 2026, headless Claude Code sessions. Output tokens include thinking; time is first to last message in each session transcript.
SettingPlanted bugs found, run 1 and 2Unplanted real bugs reportedOutput tokens, meanTime, meanList cost per run
Sonnet 5.5 low5, 50, 01,95018 s$0.07–0.08
Sonnet 5.5 medium5, 50, 02,50626 s$0.08
Sonnet 5.5 high6, 60, 03,90933 s$0.09–0.10
Opus 5.5 low6, 61, 11,81521 s$0.15
Opus 5.5 medium6, 62, 14,64252 s$0.20–0.21
Opus 5.5 high6, 61, 27,50680 s$0.25–0.28
Sonnet 5 medium4, 40, 011,419128 s$0.23–0.30
Sonnet 5 high6, 50, 018,362187 s$0.26–0.33

“Unplanted real bugs” counts how many of three genuine defects, found during the 22 September sweep but not planted, a run reported. List cost is Claude Code’s own figure for the session at API list prices. It includes the session context every run carries, about 16,000 to 17,000 tokens on the Sonnet 5.5 and Opus 5.5 runs and about 26,000 on the Sonnet 5 runs, so it overstates what the review itself cost. It is not what a subscription charges.

What it says:

What it does not say: one task, one language, one grading rule, two runs per cell. Sonnet 5’s two extra findings outside both lists were a false positive (it flagged a trailing space in a default path that is really there) and a fail-safe limitation. The same fixture flipped Opus 5.5 at low between days: on 22 September, as workflow subagents, it missed the git cherry bug in both runs, and on 29 September it caught it in both. Treat any single cell as noisy and the pattern across cells as the finding.

Effort on Sonnet 5.5#

The five levels are the same names as on Sonnet 5, and they are not the same amounts: a level on Sonnet 5.5 “doesn’t produce the same amount of thinking as the same level on Claude Sonnet 5”. Anthropic’s starting advice: “Start with high unless your workload is agentic or latency-sensitive. For agentic coding and multistep tool use, start with medium for well-specified tasks and move to high for harder or longer ones.” Effort.

My review sits on the wrong side of that line for medium. It was well specified, and medium still missed the subtle bug four times in four. For review and verification I run Sonnet 5.5 at high, which cost 33 seconds and about $0.10 a run here.

On the API, thinking is on by default and disabled returns a 400. The lowest setting is thinking: {"type": "between_tools"}, which skips up-front thinking. It is accepted at low, medium and high, returns a 400 at xhigh or max, and does not allow per-message effort changes. Migration guide.

Two behaviours from the prompting guide sit on the effort dial too:

When Sonnet 5.5 is enough#

Anthropic’s own routing starts elsewhere: “Most workloads start with Claude Opus 5.5.” Its Sonnet 5.5 launch post adds that “Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment”, and its prompting guide that “for the hardest long-horizon work, an Opus model is the better choice”. Choosing a model.

My routing, from the numbers above:

Send to Sonnet 5.5 at highKeep on Opus 5.5
The job has an acceptance check you trust: tests, a schema, a known listThe job is “find what nobody listed”: audits, reviews of unfamiliar code
High volume, where 2 times per token compoundsLong agentic runs and open-ended judgment
Scouts and extractors in a fan-outFinal verification of another model’s work
Latency matters and the task is scopedThe cheaper model would need xhigh or max to keep up

The last row is the one people miss. If Sonnet 5.5 needs xhigh or max to match, Anthropic’s own charts say Opus 5.5 at medium is often the cheaper way to the same score.

For writing and judgment I have only prior-generation data. On 22 September a decision memo scored about 34 of 40 for Opus 5.5 at every level and 26.75 for Sonnet 5, judged blind. I have not re-run it on Sonnet 5.5, so I do not carry that gap forward.

Migrating from Sonnet 5#

The model overview lists five breaking changes for code already running on Sonnet 5:

One more change fails nothing and still breaks UIs: text between tool calls now comes back in thinking blocks, so an app that streams it goes quiet between tool calls until it sets a display value that returns the text. Non-default temperature, top_p or top_k also return a 400.

Token counts match Sonnet 5, and are about 30% higher than Sonnet 4.6 for the same text. The minimum cacheable prompt drops to 512 tokens from 1,024. Images can cost more: a 2000×1500 image takes about 2.5 times the tokens it did on Sonnet 4.6. Migration guide.

Refusals, fallback and injection#

Sonnet 5.5 “declines in more categories than Claude Sonnet 5”: cyber, bio, frontier_llm, reasoning_extraction and general_harms. It is the first Sonnet to launch with classifiers against reasoning extraction. A refusal arrives as HTTP 200 with stop_reason: "refusal". A refusal that arrives before any output is billed when its category is bio, frontier_llm or reasoning_extraction, and every refusal counts against rate limits. Migration guide, refusals and fallback.

Server-side fallback is an API beta you opt into, and it retries only cyber and frontier_llm declines, on Sonnet 5. Biology requests end with a refusal and no fallback, in Claude Code too. The system card is frank about cybersecurity work: “users should expect increased refusals with Sonnet 5.5, even on benign cybersecurity-related tasks.” Finding vulnerabilities in source code is allowed; finding them in compiled binaries is blocked. Its cyber classifiers resist jailbreaks much better than Sonnet 5’s, but because it is “less cyber capable than our frontier models”, Anthropic “opted for more relaxed adversarial robustness” than on those models. System card.

Three more notes from the system card and prompting guide change how you read its output:

Price the finished job#

Per million tokensInputOutput5-min cache writeCache readBatch in / out
Sonnet 5.5$2$10$2.50$0.20$1 / $5
Opus 5.5$4$20$5$0.20$2 / $10

Cache reads cost the same on both models, so on cache-heavy agent work the per-token gap is smaller than “half” suggests. With the assumptions from Chapter 54 (a 100,000-token reusable prefix, 10,000 new input and 2,000 output tokens per request), a warm request costs $0.06 on Sonnet 5.5 against $0.10 on Opus 5.5. One cache write followed by four hits comes to $0.53 against $0.98. That is hypothetical arithmetic, not observed usage, and output tokens include thinking, which effort moves.

The measured version is the table above. On this review, Sonnet 5.5 at high cost about 36% of Opus 5.5 at high per run and 46% of Opus 5.5 at medium, with the same six bugs found and none of the extra ones. Whether that is cheaper per accepted output depends on whether your acceptance check would have caught what Sonnet missed. Divide the spend by accepted outputs, as Chapter 29 sets out, and count the reviewer’s minutes separately.

The routing rule#

Keep Opus 5.5 as the default. Move a job to Sonnet 5.5 when it has a finish line a check can verify, run it at high, and pin the full model ID. Keep the check outside the model’s own judgment, because at low it may call unverified work done. Watch the per-task cost at the top of the dial: once Sonnet 5.5 needs xhigh or max, try Opus 5.5 at medium first. For the long, open-ended, find-what-nobody-listed work, stay on Opus 5.5, and escalate beyond it only after a matched trial (Chapter 50, Chapter 49).

Use the dated tier-list reference for independent comparisons as boards add Sonnet 5.5, and the workflow planner to write the acceptance check before you choose the model. The cheaper model is the right one when your check, not the model, decides what “done” means.

Spotted something wrong, missing, or sharper? Email Vlad with feedback on this chapter →
Stay close

The next edition lands when this list says it does.

No course. No paywall. Operator playbooks weekly. 10K+ subscribers.