Claude Opus 5.5

The Effort Dial, Early Stops, and a Week of Receipts

EffortEarly stopRefusalFallbackCost per accepted output

Opus 5.5 is the model Anthropic’s docs now tell you to start with, and the one my Claude Code sessions have run since the evening of 22 September. It is cheaper than per token, it has one dial that matters, and it has two behaviours that break harnesses written for older models: it can end a turn on a progress update, and a refusal comes back as a successful HTTP response. This guide is the operating manual: what it is, what the dial buys, where it stops, what changed in the API, what a finished job costs, and what my own runs measured.

Identify the model before judging it#

Anthropic released Opus 5.5 on 22 September as the first model of the 5.5 family. The API ID is claude-opus-5-5 everywhere except Amazon Bedrock, which uses anthropic.claude-opus-5-5. The ID has no date, but it still maps to one fixed snapshot. The model takes text and images, returns text, has a 1M-token context window and a 128K maximum output (more on the Batches API with a beta header), and a June 2026 knowledge cutoff. Anthropic commits to not retiring it before 22 September 2027 on its own platforms. Model overview, model IDs.

Where it runs matters as much as what it is:

What Anthropic claims, and how to read it#

The launch post’s headline is that Opus 5.5 “performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5”. Both halves need their conditions. The 40% is Anthropic’s test result “at default settings” on “typical workloads”, not a price cut. The price cut is 20% on input and output, to $4 and $20 per million tokens, and 60% on cache reads, to $0.20. The same post claims output “more than 30% faster than Opus 5” and gives no method for it. Launch post, what a task costs.

The launch table below keeps each score’s conditions. Unless the conditions column says otherwise, Opus 5.5 ran at max effort, averaged over five trials, with safeguards on. Safeguards on means a fallback model answered some items: on Terminal-Bench 4.0 the fallback served 2.5% of requests, touching 10% of trials, so a model column is not a pure-model run. System card, p. 178.

BenchmarkOpus 5.5Fable 5.1Opus 5GPT-6 AstraConditions
Terminal-Bench 4.066.4%55.8%52.3%57.9%Opus 5.5 at xhigh (64.8% at max); Astra at high, reported by OpenAI
FrontierCode v1.154.4%50.3%48.0%53.3%Run by Cognition. At each model’s best effort: 54.6% against Opus 5’s 53.4%
GDPval-AA v2.11846173517081542Elo, run by Artificial Analysis. Opus 5.5 at medium: 1576
SWE-bench Pro89.9%81.2%79.2%—System card only
AutomationBench40.0%31.4%26.9%41.4%Run by Zapier without fallback, so safeguard stops counted as failures (42.5% with fallback in the Sonnet 5.5 card); Astra leads
FrontierSWE v262.3%56.3%—65.5%Proximal’s harness; Astra leads
Toolathlon Verified77.8%77.8%80.6%—Pass@1 over three trials, seven stopped trials counted as failures; below Opus 5

Two readings follow. First, “leads in agentic coding” is true of Anthropic’s chosen boards, not all of them: in Anthropic’s own tables Astra is ahead on AutomationBench, FrontierSWE v2 and Terminal-Bench-Science, and Toolathlon went backwards from Opus 5. Second, the FrontierCode lead over Opus 5 is 6.4 points at max and 1.2 points at each model’s best setting, because Opus 5.5’s score peaks at medium. The post itself adds the caveat worth keeping: “at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.”

What independent boards found#

These were captured on 29 September, a week after launch. Every board runs its own harness, so the numbers match neither Anthropic’s nor each other’s. That mismatch is the first finding.

Across the boards, the gains live at medium and high. max buys little and can cost more.

Effort is the only dial#

On Opus 5.5 thinking is always on. A request that disables it, or sets budget_tokens, returns a 400. What you control is : low, medium, high, xhigh, max. The API default is medium, while every other model that supports effort defaults to high, so a request that omits the parameter runs one level lower than it did on Opus 5. Effort.

Level names do not carry across models. At a given level Opus 5.5 “tends to think more per turn than Claude Opus 5, most of all at xhigh and max”, and Anthropic’s testing puts its medium level with or above Opus 5 at high. Carry an old high over and you buy longer turns. What’s new, prompting guide.

Anthropic’s own cost numbers make the trade concrete. On a 478-problem internal SWE-bench Pro subset (not comparable to the public leaderboard), measured against high, medium scored about 2.5 points lower for about 70% of the cost, low about 8 points lower for about a third, and xhigh about 1.4 points higher for 2.5 times the cost. The same page describes the cheaper policy: run everything at low, then re-run only the failures at high. That passed about 97% of tasks for about $0.17 each, against 95.3% for $0.29 running everything at high, counting the failed cheap attempts. Anthropic’s own advice is to use it “for the saving, not the lift”. Optimizing for cost and intelligence.

The Claude Code team’s effort post maps the levels to work. low is for quick in-the-loop work like brainstorming, sketching and easy changes. medium is for most regular engineering, such as implementing a new feature. high is for work where verification matters or there are edge cases, like a bug fix in a brownfield codebase. max is for fully autonomous hard problems, like building and verifying an app end to end. Its reading of Terminal-Bench 3.0 is that higher effort pays most on tasks with hidden edge cases. On that board Opus 5.5 went from 36.6% at low to 65.7% at max, and matched Fable 5.1 at max from high (58.9% against 58.0%) on half the tokens.

Three mechanics bite in practice:

What I measured#

On 22 September, the night Opus 5.5 became my main model, I swept it across effort levels on three tasks, two runs per setting. The discriminating one was a code review. The file was a 209-line Python script that retires git worktrees, with six bugs planted in it. The task was to report every place where the code breaks its own docstring or could lose data. Code graded the reports: a finding within two lines of a planted bug counts as found. Opus 5 and Sonnet 5 at high ran as references. The runs were Claude Code workflow subagents.

One review, four effort levels: output tokens, time and bugs found
One review, four effort levels: output tokens, time and bugs found Six planted bugs, two runs per setting, 22 Sep 2026, Claude Code workflow subagents. Output tokens include thinking; time is first to last message in each agent transcript.
SettingPlanted bugs found, run 1 and 2Output tokens, meanTime, mean
Opus 5.5 low5, 51,50718 s
Opus 5.5 medium6, 64,05443 s
Opus 5.5 high6, 68,56487 s
Opus 5.5 xhigh6, 619,268185 s
Opus 5 high6, 612,326152 s
Sonnet 5 high6, 416,815564 s

medium found everything with less than half the output of high. xhigh used 2.25 times the output tokens and 2.1 times the time of high, and found no more. low missed the same bug in both runs, the subtlest of the six: a git cherry check that reads the wrong marker and so calls unmerged work merged. Opus 5.5 at high matched Opus 5 at high on the bugs with about 30% fewer output tokens. The other two tasks did not separate the levels. Extraction, twelve questions against a 281 KB document, scored 12 of 12 on every setting. A decision memo capped at 450 words, scored blind out of 40 by a Sonnet 5 judge and an Opus 5 judge, averaged between 33.75 and 34.75 at every Opus 5.5 level and 34.75 for Opus 5, against 26.75 for Sonnet 5.

On 29 September I re-ran the review as headless sessions (claude -p) with the same prompt, beside Sonnet 5.5 (see Chapter 55). Opus 5.5 found all six bugs in both runs at every level, low included. The low result flipped between days and harnesses, which is the honest size of a two-run sample. Two things held. Output grew with effort (about 1,800, 4,600 and 7,500 tokens at low, medium and high), and every Opus 5.5 run also reported at least one of three real bugs nobody planted, which no Sonnet run did.

What I run now, from those numbers: extraction and loaders at low, scouts at medium, and finders, builders and verifiers at high. My saved session level is high too. xhigh and max are reserved for work where I have measured a gain. On this review, medium would have been enough. I pay for high in review stages because one of its runs reported all three real bugs nobody planted, where medium reported two, and a missed bug costs more than the 44 seconds high adds.

A week of daily use adds one more receipt. From the evening of 22 September to early afternoon on 29 September, my local Claude Code transcripts hold 53,445 assistant messages served by Opus 5.5, across 976 transcript files, counted once per message ID. Four ended in a refusal. All four were reasoning_extraction, all in one session of one project, at xhigh. They came after short operational requests, one of them a five-word request to confirm it was working for users, and none asked the model to reveal its reasoning. Claude Code posted its “stopped by a safety classifier” notice, and each time the next turn carried on with Opus 5.5 within seconds. That is about one refusal per 13,000 messages, clustered rather than random.

Long runs can stop on a progress update#

On long tasks with several parts, Opus 5.5 keeps the user updated as it works, and “some of those updates end the turn with text rather than a tool call”. A harness that reads a text-only end of turn as “done” stops there, with the work unfinished and a clean exit code. Prompting Opus 5.5.

Anthropic’s fixes are harness fixes:

The same behaviour changes streaming UIs. On the API, text Opus 5.5 writes between tool calls now arrives as thinking blocks, empty under the default display, so an app that streamed that text as progress “goes quiet between tool calls”. Use the display: "updates" beta, or a send-message tool declared from the first request. Migration guide.

Judge an unattended run by its artifact, never by its exit code. Chapter 38 has the harness side of this.

Refusals arrive as HTTP 200#

A declined request returns HTTP 200 with stop_reason: "refusal" and a stop_details.category: cyber, bio, frontier_llm, reasoning_extraction or general_harms. Monitoring that only watches for errors will count a refusal as a success. Since 24 September, a refusal that arrives before any output is billed when its category is bio, frontier_llm or reasoning_extraction, and every refusal counts against rate limits. What’s new, release notes.

On the API, server-side fallback is a beta you opt into. It is not available on Bedrock, Google Cloud, Foundry or the Batches API, and it never retries a reasoning_extraction refusal. The system card names the targets: cyber flags go to Opus 4.8, and biology and frontier-LLM flags go to Opus 5. Distillation and weapons requests are blocked with no fallback. In the Claude apps the switch is automatic, and the checks review everything the model reads, including memory, connector content, web search results and files, so content you did not type can trigger a fallback. Why Claude switched models.

The cyber policy allows vulnerability-finding in source code and blocks it in compiled binaries, with what the system card calls “a temporarily wider safety margin”. Opus 5.5 is not yet in the Cyber Verification Program. Security teams should expect friction: on Artificial Analysis’s CyberGym-E2E-AA, where the model must prove a memory-safety crash and then patch it, Opus 5.5 refused at least 98% of the tasks, as did GPT-6 Astra and Fable 5.1. Artificial Analysis cyber index. The prompting guide warns that a prompt pushing the model to reproduce its reasoning in the response can be declined as reasoning_extraction. Its fix is to remove such instructions and read the reasoning from summarized thinking blocks (display: "summarized") instead. My four refusals above followed prompts that asked nothing of the kind, so expect the occasional false positive anyway.

Migrate the contract, not the string#

Changing the model ID is the easy line of the diff. Check these before you move traffic, against the migration guide:

Price the finished job#

Per million tokensInputOutput5-min cache write1-hour cache writeCache readBatch in / out
Opus 5.5$4$20$5$8$0.20$2 / $10
Sonnet 5.5$2$10$2.50$4$0.20$1 / $5
Fable 5.1$10$50$12.50$20$0.25$5 / $25

There is no surcharge for long context, and US-only inference costs 1.1 times list. Pricing, what’s new.

Hypothetical arithmetic, not observed usage, with the same assumptions as the Fable 5.1 chapter: a 100,000-token reusable prefix, 10,000 new input tokens and 2,000 output tokens per request, Standard service, no tools.

ConditionCalculation in USDTotal
Cold, no cache0.11 × $4 + 0.002 × $20$0.48
Cold, five-minute cache write0.10 × $5 + 0.01 × $4 + 0.002 × $20$0.58
Warm, matching prefix hit0.10 × $0.20 + 0.01 × $4 + 0.002 × $20$0.10

One write followed by four hits comes to $0.98, against $2.40 for five uncached requests. The same pattern costs $2.35 on Fable 5.1. Output tokens include thinking, so the effort level moves the output line. My measured review runs on 29 September cost about $0.15 at low, $0.21 at medium and $0.26 at high at list price, as Claude Code reported them. Each figure includes the roughly 17,000 tokens of session context every headless run carries.

Two traps sit outside the token table:

Divide all run spend by accepted outputs, as Chapter 29 sets out. A run with no accepted output cost money; it did not cost zero.

Prompts: what to strip, what to add#

Strip:

Add, where they fit:

All quotes are from Prompting Opus 5.5 and the system card.

What the system card adds#

The system card is 230 pages. Five findings change how you operate the model:

The routing rule#

Start on Opus 5.5 at medium for regular engineering, and raise it to high for review, verification and bug fixes in code you did not write. Reach for xhigh or max only after a measured gain on your own task, and re-run failures at a higher level rather than running everything high. Move well-scoped work with a checkable finish to Sonnet 5.5 at high (see Chapter 55 for the same review run on both). Escalate to Fable 5.1 or GPT-6 Astra only after a matched trial on your own task (Chapter 50, Chapter 49).

If your main loop and your subagents both run Opus 5.5, an unpinned judge is the same model grading itself. Pin verification stages to a different model, with a fallback chain for when that model’s limit runs out. Keep the dated tier-list reference for independent comparisons, and turn one candidate task into a specification with the workflow planner before you change a default. Pin the model ID, log the model that served each step and the effort it ran at, and count every attempt. After that, the upgrade is a routing rule you can defend.

Spotted something wrong, missing, or sharper? Email Vlad with feedback on this chapter →
Stay close

The next edition lands when this list says it does.

No course. No paywall. Operator playbooks weekly. 10K+ subscribers.