Model file · Opus 5

Opus 5 use cases — seven jobs, and the dial setting for each.

Most launch coverage answers "is it better." That is the wrong question, because the honest answer is statistically indistinguishable from the model above it, at half the price. The useful question is which jobs, at which effort — because on this model the dial is worth more than the model choice, and on one job below turning it up makes the output worse.

Ranked by strength of evidence, not by excitement. Every row names its source and its trust tier. The specs, pricing and the effort ladder live on the model file; the discipline for reading any of these numbers is Ch 24.

Jump to section tap to open

The 30-second answer

Opus 5 is a long-horizon agent model. Its best-evidenced win: on AA-Briefcase, at the default effort setting, it scores higher than Fable 5 at max for 47% of the cost. Use high for long-horizon work, medium for automation and in-repo debugging (where scores peak), low for subagents, max only for computer use. It is the wrong tool for anything latency-sensitive, for high-volume repeated work, and for review tasks where recall beats precision.

The seven jobs, ranked by evidence

1

Long-horizon agentic knowledge work

Effort: high — escalate to xhigh only past 30 minutes

The strongest evidence in the entire release, and the one that should change a default today.

On AA-Briefcase — Artificial Analysis's long-horizon knowledge-work benchmark, built around multi-week projects — Opus 5 at the default high scores 1606 Elo for $10.41 a task, while Fable 5 at max scores 1574 for $22.30. A higher score at 47% of the cost, from the tier below, at the setting you get for free by not touching anything. Decomposed, the gap is lopsided: Analytical Quality Elo 2016 against Fable 5's ~1716, while Presentation Quality actually goes to GPT-5.6 Sol. Opus 5 thinks better than it writes.

Turning it up: max buys +114 Elo for +71% cost, +10 minutes and +27 turns per task. xhigh is the honest ceiling at 1693 for $14.26 — and Anthropic reserves it for work "over 30 minutes with token budgets in the millions."

The caveat: Opus 5 runs ~50% longer and takes nearly 2× the turns of Opus 4.8 on identical tasks — 36.2 minutes and 103 turns against 24.1 and 55. The quality is bought with sustained loop length, which makes cost forecasting harder, not easier.

Artificial Analysis · 2026-07-24 · independent

2

Workflow and SaaS-integration automation

Effort: medium

The best cost-per-outcome row in the launch, and almost nobody quoted it.

Zapier ran Opus 5 against AutomationBench — their own harness across 47 simulated SaaS tools, scored on a held-out split. On Zapier's leaderboard it takes 23.6% at $0.89 a task at medium effort, beating both Opus 4.8 (17.0) and Fable 5 (17.4) at less than half their cost. At max it reaches 26.2% for $1.27. This is an independent leaderboard from a company with no incentive to flatter, which puts it a full trust tier above the launch table.

Turning it up: Barely worth it — max buys 2.6 points for 43% more spend. The medium-effort row is the headline here.

The caveat: Only the leaderboard's scoring split is held out: Zapier open-sourced a 600-task public set across the same tools and reports local runs in directional agreement with the board, so this is one of the few rows on this page you can actually re-run yourself. Anthropic's own summary table prints slightly different figures (24.0 and 26.0); Zapier's are used here, because it is Zapier's benchmark.

Zapier AutomationBench leaderboard · 2026-07-24 · INDEPENDENT

3

Debugging and root-cause analysis in an existing repo

Effort: medium — do not go above high

The one job where turning the dial up makes the output measurably worse.

Cognition ran and scored FrontierCode 1.1: Opus 5 hits 53.4 on Main, and its best score is at medium. Anthropic documents the decline itself — "a decline in FrontierCode score above high effort… Opus 5 at these effort levels [makes] more changes than the task requires (e.g., refactoring)." Devin's qualitative read matches: "particular strength on difficult debugging and root-cause analysis," "prefers targeted, in-place fixes over large refactors," "adhering to existing repo conventions."

Turning it up: Escalating is the mistake. Above high you are paying more to get scope creep.

The caveat: Cognition is a launch partner — independent harness, but not a disinterested party. That said, the vendor and the partner and the practitioners all report the same decline, which is unusual and worth weighting.

Cognition · System Card §8.4 · 2026-07-24 · partner-independent

4

Computer use and desktop automation

Effort: max

The largest single jump over Opus 4.8 anywhere in the card — and the most misreadable number in the launch.

OSWorld 2.0: 70.6 against Opus 4.8's 55.7, Fable 5's 66.1 and GPT-5.6 Sol's 62.6. A +14.9-point jump over the previous Opus in one release.

Turning it up: This is the one row where max is the documented setting and the gain justifies it.

The caveat: Two things must travel with this number or it lies. It is partial/checkpoint credit, not task completion — the OSWorld 2.0 paper scores Opus 4.8 at 20.6% binary against 54.8% partial, so "70.6% on computer tasks" reads roughly 3× what it means. And Opus 5 is absent from the official OSWorld 2.0 leaderboard, which still tops out at Opus 4.8. Vendor-only, with Opus 4.8 as the grader.

Anthropic System Card Table 8.1.A · 2026-07-24 · vendor

5

Novel-abstraction and open-ended puzzle reasoning

Effort: high (max was never tested)

The one launch-window number this site credits without a discount — because a third party ran it.

ARC-AGI-3, Public Demo set: 30.16%, against GPT-5.6 Sol's 13.33% on the same set and Anthropic's Fable-class models at ~20%. Opus 5 "completed five additional Public Demo environments that no model had previously beaten." ARC Prize's judge on the game Axis Reflect: Opus 5 scored 100, cleared all eight levels in 294 actions, derived an explicit reflection equation by level 2 and generalized to 2D by level 8 — "once the correct ontology is found, execution is extremely reliable."

Turning it up: Not measurable. ARC Prize evaluated high only, "due to the short testing window" — no max-effort ARC-AGI-3 figure for Opus 5 exists.

The caveat: Anthropic's own table compares this against a GPT-5.6 Semi-Private score, which mixes evaluation sets and inflates the gap from ~2.3× to ~4×. The comparison above is same-set.

ARC Prize Foundation · arcprize.org · 2026-07-24 · INDEPENDENT

6

Single-window long-context work

Effort: high, with a fresh 1M budget per pass

Where the pricing structure, not the benchmark, is the actual advantage.

ProgramBench — 166 golden tasks, 247,000+ behavioral tests, a fresh 1M-token context per episode: 83% after episode 1, rising to 93% by episode 5. Beats Opus 4.8 (80→90) and ties Mythos 5 (84→93). Pair that with the pricing fact from the model file: there is no long-context surcharge, so a 900K-token request bills at the same per-token rate as a 9K one, while GPT-5.6 charges 2× input and 1.5× output above 272K.

Turning it up: Unnecessary — the win here is structural, not dial-dependent.

The caveat: Anthropic's claim that "instruction following, tool calling, and reasoning stay consistent throughout the window" is asserted, not independently benchmarked. Nobody has published a needle-in-a-haystack-style degradation curve for Opus 5.

Anthropic System Card · 2026-07-24 · vendor

7

Agentic search and multi-hop research

Effort: high — and bring your own fetcher

Strong numbers with a hard architectural blocker attached.

BrowseComp 90.8 (GPT-5.6 Sol 90.4, Fable 5 87.4) and Humanity's Last Exam with tools at 64.7. A 10-agent team reached 93.6% with a 5.9× latency speedup at N=10.

Turning it up: The multi-agent figure was gathered on a pre-release configuration of Opus 5 — treat it as directional only.

The caveat: The web_fetch server tool is not available on Opus 5. If your search agent depends on Anthropic's server-side fetcher, Opus 5 cannot run it as written — you must supply your own. Opus 5 also has no entry on Arena's Search board at all.

Anthropic System Card · 2026-07-24 · vendor

Eleven places Opus 5 is the wrong call

A model file that only lists strengths is a brochure. These are the documented cases where something else wins — several of them from Anthropic's own guidance, one of them a production A/B that contradicts the launch narrative outright.

01 Anything latency-sensitive

Time-to-first-token is 21.7s at the default high and 66.4s at max, with throughput flat at 52–56 tok/s against a ~75 t/s median. Artificial Analysis ranks it #117 of 190 on speed and calls it "notably slow and very verbose." Anthropic's own model matrix routes real-time applications to Haiku 4.5. Fast mode buys up to 2.5× output speed at $10/$50 — Fable 5 token pricing for an Opus-tier model, and it is a research preview.

02 As the leaf node in a subagent swarm

Anthropic's effort table names low as the subagent setting and its model matrix names Haiku 4.5 for sub-agent tasks. Running Opus 5 at high in every leaf of a depth-3 tree is the documented anti-pattern — and depth-3 became the default in Claude Code v2.1.219, shipped the same day as the model. Practitioners report subagent eagerness multiplying cost by "an order of magnitude."

03 High-volume, well-defined, repeated work

Sonnet 5 costs $1.53 per Index task to Opus 5's $2.03, at $2/$10 through August 31 (then $3/$15) against $5/$25. Anthropic routes high-volume intelligent processing to Haiku 4.5 and scaled code generation and data analysis to Sonnet 5.

04 Code review where recall matters more than precision

CodeRabbit's production A/B is the sharpest counterexample to the launch narrative. Opus 5 at xhigh caught 55.2% of known issues against CodeRabbit's production model mix at 61.1% — and GPT-5.6 Sol on its own catches 69.7%. It generated 92 nitpicks against 23 and burned ~50% more input and ~65% more output tokens per call. On the comments it marks actionable it is more precise (39.3% vs 35.2%), but across the full post-pipeline stream it is less precise (28.6% vs 32.8%) — so "a precision play, not a coverage play" holds only for the filtered slice. CodeRabbit's own caution: small numeric differences should not be read as head-to-head wins. Anthropic's own guidance if you use it here: do not tell it to be conservative, because "the model may follow that instruction literally and report less." Ask for everything and filter in a second pass.

05 When you need the vendor-declared ceiling

Anthropic's matrix still assigns "the highest available capability / long-running agents / advanced research" to Fable 5. The official tbench.ai Terminal-Bench 2.1 #1 is Claude Code + Fable 5 at 83.8%, and Opus 5 has no submission at all.

06 Knowledge-reliability-critical work

On AA-Omniscience, accuracy rises +7 points over Opus 4.8 — but hallucination rate rises +14 points, to 50%. The composite index tells the story: Fable 5 40.15 against Opus 5's 31.27. Pair with the HealthBench collapse from 67.1 raw to 57.8 length-adjusted. Fable 5 is materially more reliable on knowledge.

07 Offensive security and exploit development

Anthropic states Opus 5 is "substantially weaker on exploits" than Mythos 5, and binary/compiled-code vulnerability discovery remains blocked. Source-code vulnerability discovery is now permitted at all access levels. One published long-horizon signal: Zvi Mowshowitz reports Opus 5 approaching Mythos 5 on ExploitBench at a 2-hour budget but falling behind at 6 — in a domain Anthropic deliberately suppressed, so it does not generalize.

08 Integrations that must run with thinking disabled

A documented failure mode: Opus 5 "occasionally writes a tool call into its user-facing text instead of emitting a structured tool_use block. The turn completes normally and the call never runs, and in agentic loops the leaked text stays in the conversation history." Most common on tool-heavy search workloads. Anthropic's fix: "for most tasks, thinking enabled at low effort performs better than thinking disabled at similar cost."

09 Structured or templated output where verbosity is a cost

Default responses, agentic narration and written files all run longer than Opus 4.8 — and effort will not fix it. Anthropic, verbatim: "changing effort does not reliably shorten responses, so prompt for length instead."

10 Long cached sessions with dynamic effort escalation

Changing effort mid-conversation invalidates the prompt cache, because effort shapes the rendered prompt. Pick one level at the start and keep it.

11 Workloads needing Priority Tier or server-side web_fetch

Neither is supported on Opus 5. Both work on Opus 4.8. This is the rare case where the newer model is the downgrade.

What practitioners say

Praise
  • Simon Willison flags the example where the model "independently developed a computer vision pipeline to analyze machine parts when direct viewing wasn't available" — while explicitly noting he had not yet put it through its paces. The widely-quoted "our least prompt injectable model yet" line is Boris Cherny's, an Anthropic employee — a vendor claim, not third-party praise. Willison relayed it without endorsing it, and it does not belong in this column.
  • Zvi Mowshowitz on the system card: prompt-injection attack success drops 5.5% → 2.0%; computer-use environment attack rate 7.14% → 0.54%. When the model found scoring exploits it declined, calling them "hacks."
  • Cognition's CEO: strongest on difficult debugging and root-cause analysis inside Devin.
  • Hacker News on the economics: "Half the price of Fable 5 and usable with 100% of your subscription means roughly 4× the usage." Also valued: mid-conversation tool swapping, automatic Opus 4.8 fallback, and no 30-day retention requirement.
Complaints
  • Scope creep above high — "Above high, Opus 5 starts making changes the task didn't ask for." This is the most load-bearing community claim in the set, because it is corroborated by two non-community sources: Anthropic's own scope-expansion warning and Cognition's measured decline.
  • Token burn with no pushback — "it will almost certainly not push back on your silly request and go ahead and burn as many tokens as it can." Recurring calls for per-prompt token caps and an explicit autonomous-vs-interactive mode.
  • Over-delegation — subagent eagerness reported to multiply cost by "an order of magnitude."
  • Truncation — thinking counts against max_tokens; migrated code cuts off mid-sentence.
  • Regression on clarification — Opus 4.8 was "generally better at asking for clarification."
  • Benchmark-integrity criticism — the undisclosed Opus-4.8-fallback substitution on FrontierBench, and several benchmarks built by outside orgs but run by Anthropic.

FAQ

What is Claude Opus 5 best at?

Long-horizon agentic knowledge work is the strongest evidenced use case: on Artificial Analysis's AA-Briefcase benchmark, Opus 5 at its default effort setting scores higher than Claude Fable 5 at maximum effort while costing roughly half as much per task. Beyond that, the evidence supports workflow automation at medium effort, in-repo debugging and root-cause analysis at medium effort, desktop automation at max effort, and novel-environment reasoning — where its ARC-AGI-3 result was administered by the ARC Prize Foundation rather than by Anthropic.

What effort level should I use with Claude Opus 5?

Start at high, which is the API default and equivalent to not setting the parameter. Use medium for workflow automation and for debugging inside an existing repository, where scores actually peak. Use low for subagents. Reserve xhigh for genuinely long-running work past thirty minutes, and max only for computer use. Do not inherit an xhigh default from Opus 4.7 or 4.8 — Anthropic explicitly recommends running a fresh effort sweep instead.

When should I not use Claude Opus 5?

Avoid it for anything latency-sensitive — time to first token is about 21.7 seconds at the default setting and 66.4 seconds at maximum. Avoid it as the leaf node of a subagent swarm, where Anthropic's own guidance names Haiku 4.5. Avoid it for high-volume repeated work, where Sonnet 5 is cheaper per task. And avoid it for code review where recall matters more than precision: one production A/B by CodeRabbit found it caught 55.2% of known issues against 61.1% for their production model mix, while producing four times the nitpicks.

Does Claude Opus 5 hallucinate more than Claude Fable 5?

On Artificial Analysis's AA-Omniscience benchmark, yes. Opus 5 gains about 7 points of accuracy over Opus 4.8 but its hallucination rate rises about 14 points, to roughly 50%. The composite AA-Omniscience index puts Claude Fable 5 at 40.15 against Opus 5 at 31.27. For knowledge-reliability-critical work, Fable 5 is the more reliable choice despite the near-identical aggregate intelligence scores.

Keep reading

The Opus 5 model file — specs, pricing, the full effort ladder, and where it sits against Fable 5. Ch 24, the tier list — the four boards it now appears on, and why they disagree. Ch 25 — how to run the effort sweep this page keeps telling you to run.

Stay close

The next edition lands when this list says it does.

No course. No paywall. Operator playbooks weekly. 10K+ subscribers.