The tier list. Without mercy.
Public leaderboards rank capability. Operators rank usefulness. Below, five readings of the same models: the crowd's public benchmark, an independent referee that prices every run, the labs' own launch decks, the one crowd board that scores agents doing work — then Vlad's actual stack, a builder you can drag, share, and argue about.
Two Anthropic models shipped since the last capture — Sonnet 5 on June 30 and Opus 5 on July 24 — and one of them broke the rule this page was built on. For as long as these boards have existed, the most capable model was also the most expensive one. It isn't any more: Opus 5 took the top of the independent index at lower cost per task than the model below it, and it ships with a five-position effort dial that moves capability and the invoice together. The Opus 5 model file has the full read, and the use-case page has the part that changes what you type tomorrow.
Also new: LMArena is now Arena, its boards no longer share an update date (the spread in this capture is 34 days, so each carries its own), and its HuggingFace feed was retired — so the leaderboard below is a hand-verified mirror, cross-checked against a community mirror on ten of eleven boards.
Arena public leaderboard
Crowdsourced head-to-head model votes from arena.ai, the platform formerly called LMArena. This is the canonical "which model is better at benchmark tasks" public ranking. It's the right answer to one specific question — and the wrong answer to the question you actually care about as an operator.
Two things to hold while reading it. First, the sample sizes are not comparable across rows: Opus 5 debuts at #1 on Document on 1,663 votes against the 37,271 behind the model it displaced, with fully overlapping confidence intervals. That is a leaderboard fact, not yet a statistical one, and the per-board notes say so where it applies. Second, Opus 5 appears here as claude-opus-5-high — the crowd is voting on the API default effort, while every vendor table further down runs at max. Same model, different configuration, different number.
Artificial Analysis — the independent referee
The arena above is the crowd's taste; the boards below are the vendors' launch slides. This is neither — Artificial Analysis runs its own agentic harness (rebuilt in Intelligence Index v4.1 around long agentic chains — Terminal-Bench 2.1, 𝜏³-Banking, GDPval v2) and reports the one number the other two never do: what each run costs. Cost per task, not just capability — the column where the leaderboard starts speaking the operator's language. Read it as a referee's reading, not a verdict: independent buys disinterest, not the last word. Captured 2026-07-27; dated sourcing on the research notes timeline.
Two disclosures this board earns. AA states it "supported Anthropic to evaluate Claude Opus 5 ahead of release" — its own harness and hardware, so not a vendor claim, but pre-release coordination on the row that came first. And the top of this board is a tie, not a ranking: AA's own stated confidence interval is ±1%, and the gap between the top two distinct models is 0.83 points. (AA ranks effort variants separately, so its literal second place is the same model again at a lower setting, 0.62 behind.) A page that discounts launch decks has to apply the same arithmetic to the referee.
How the Index is built — v4.1 (June 2026 — current)
- Agents · 34% — GDPval-AA v2 · 20% · 𝜏³-Banking · 14%
- Coding · 24% — Terminal-Bench v2.1 · 16% · SciCode · 8%
- Scientific reasoning · 24% — HLE · 12% · GPQA Diamond · 6% · CritPt · 6%
- General · 18% — AA-Omniscience · 12% (accuracy 8% + non-hallucination 4%) · AA-LCR · 6%
Launch-deck numbers — discounted on arrival
The arena above is crowd votes on chat prompts. The boards below are what vendors shipped on launch day for the question operators actually ask — agentic work — which the arena doesn't measure. Read every row as a claim, not a receipt: Berkeley RDI reward-hacked 8 of 8 major agent benchmarks, so launch numbers carry a 10–15 point discount until your own eval confirms them. Models you can't buy don't appear — capability ceilings are not tier-list entries. Dated sourcing for every figure lives on the research notes timeline.
Three of these four cards are vendor claims. One was run by a third party. Watch what happens to the numbers when the referee changes — that difference is the entire argument of this chapter, and it is now visible in a column rather than asserted in a sentence.
Max effort, mean of 5 trials — max is NOT the API default. Fable 5 still wins this row. Opus 5 appears on no independent SWE-Bench board. And Cursor's June study found 63% of Opus 4.8's passes retrieved a known fix rather than deriving one; sealing git history AND cutting internet access together dropped it 87.1% → 73.0% (of audited retrievals, web lookup was 57% and git-history mining only 9% — the internet is the bigger leak). Discount accordingly.
The most instructive row on this page, and not for the reason you would expect: the independent board AGREES on Opus 5 (43.5 ±1.65, rank 1) and disagrees sharply about its rival — Anthropic re-ran GPT-5.6 Sol itself at 37.5, while the public board has Sol at 34.4 on OpenAI's own agent. Two caveats survive the agreement. Anthropic's separate §8.5 run scores best at xhigh (44.4), not the max its headline prints. And in that run Opus 4.8 silently substituted whenever a safety classifier refused — 4% of Opus 5's trials, 26% of Fable 5's — so both Anthropic scores are two-model blends while the OpenAI score is clean. Note also that the public board mixes harnesses: Opus 5's rank-1 row runs on Princeton's mini-SWE-agent, Fable 5's on Anthropic's own.
This is PARTIAL/checkpoint credit, not task completion — the OSWorld 2.0 paper scores Opus 4.8 at 20.6% binary against 54.8% partial, so "70.6% on computer tasks" reads about 3× what it means. Opus 4.8 was the grader. The GPT-5.6 figure was lifted from OpenAI's own post, so the gap is cross-harness. Opus 5 is absent from the official board.
The one launch-window number on this page that earns full credit, because a third party executed and scored it. Read the config honestly: ARC Prize evaluated HIGH effort only — "due to the short testing window" — so no max-effort figure exists. All three rows above are the same Public Demo set; Anthropic's own table mixes in a Semi-Private score for GPT-5.6 (7.8%), which inflates the gap from ~2.3× to ~4×. One unresolved discrepancy worth stating rather than smoothing: Anthropic's §8.14.2 describes these same results as semi-private, while ARC Prize's own results page files them under the Public Demo set. We follow ARC Prize, since ARC Prize ran it.
Arena's Agent board — and why it isn't a tab above
Arena runs an agent board too, over 1,242,857 sessions · 38 models. It is deliberately not one of the tabs above, because it does not use Elo — its metric is a Net Improvement Score, a percentage. Rendering 12.72 in a column that reads 1508 two tabs over would be the exact category error this chapter exists to argue against, so it gets its own shape. It is also the closest any crowd board comes to measuring what operators actually buy — and note who is missing from it.
Vlad's tier list
The arena measures capability head-to-head on neutral prompts. This list measures what runs Vlad's portfolio on a Tuesday at 11am. They disagree more than you'd expect — Claude Code and Cowork don't appear on the arena at all, Perplexity loses head-to-head but is S-tier here because it answers the actual question 90% of the time. Drag tools, build your own, share the URL.
Benchmarks reward the model that beats other models. Operators reward the tool that doesn't break the workflow. The most useful tool you own is rarely the most capable one — it's the one with the lowest activation energy on a Tuesday morning when you have nineteen other problems. That's why Claude Code is S-tier in this list and not even ranked on LMArena. Different question. Different answer.