Chapter 24

The tier list. Without mercy.

Public leaderboards rank capability. Operators rank usefulness. Below, five readings of the same models: the crowd's public benchmark, an independent referee that prices every run, the labs' own launch decks, the one crowd board that scores agents doing work — then Vlad's actual stack, a builder you can drag, share, and argue about.

Refreshed 2026-07-27

Two Anthropic models shipped since the last capture — Sonnet 5 on June 30 and Opus 5 on July 24 — and one of them broke the rule this page was built on. For as long as these boards have existed, the most capable model was also the most expensive one. It isn't any more: Opus 5 took the top of the independent index at lower cost per task than the model below it, and it ships with a five-position effort dial that moves capability and the invoice together. The Opus 5 model file has the full read, and the use-case page has the part that changes what you type tomorrow.

Also new: LMArena is now Arena, its boards no longer share an update date (the spread in this capture is 34 days, so each carries its own), and its HuggingFace feed was retired — so the leaderboard below is a hand-verified mirror, cross-checked against a community mirror on ten of eleven boards.

What the public says

Arena public leaderboard

Crowdsourced head-to-head model votes from arena.ai, the platform formerly called LMArena. This is the canonical "which model is better at benchmark tasks" public ranking. It's the right answer to one specific question — and the wrong answer to the question you actually care about as an operator.

Two things to hold while reading it. First, the sample sizes are not comparable across rows: Opus 5 debuts at #1 on Document on 1,663 votes against the 37,271 behind the model it displaced, with fully overlapping confidence intervals. That is a leaderboard fact, not yet a statistical one, and the per-board notes say so where it applies. Second, Opus 5 appears here as claude-opus-5-high — the crowd is voting on the API default effort, while every vendor table further down runs at max. Same model, different configuration, different number.

Arena · Text
The headline board — head-to-head chat votes.
captured 2026-07-27
board published 2026-07-26 · 7.48M votes
1
Anthropic
claude-fable-5
1508
2
Anthropic
claude-opus-4-6-thinking
1505
3
Anthropic
claude-opus-4-7-thinking
1502
4
Anthropic
claude-opus-4-6
1498
5
Anthropic
claude-opus-5-high
1495
6
Anthropic
claude-opus-4-7
1493
7
Meta
muse-spark-1.1
1493
8
Meta
muse-spark
1488
9
Google
gemini-3.1-pro-preview
1486
10
Google
gemini-3-pro
1486
11
Moonshot
kimi-k3
1485
12
OpenAI
gpt-5.6-sol-xhigh
1485
Anthropic holds 6 of the top 10 here. The Elo spread across this top 10 is 23 points — a gap operators rarely feel in practice.
Read the board, not the rank. Anthropic holds all six of the top six — a clean sweep — and no OpenAI model reaches the top 10; the highest GPT entry is #12. Opus 5 debuts 5th, behind three older Opus variants, on 5,417 votes against 16k–68k for the models above it (±8 vs ±4–6). Ranks 9–12 span a single Elo point.
Crowdsourced head-to-head votes, hand-mirrored on 2026-07-27. Arena ships no public API; ten of eleven boards were cross-checked against a community mirror — a corroborating source, never the citation.Open arena.ai →
What an independent referee measures

Artificial Analysis — the independent referee

The arena above is the crowd's taste; the boards below are the vendors' launch slides. This is neither — Artificial Analysis runs its own agentic harness (rebuilt in Intelligence Index v4.1 around long agentic chains — Terminal-Bench 2.1, 𝜏³-Banking, GDPval v2) and reports the one number the other two never do: what each run costs. Cost per task, not just capability — the column where the leaderboard starts speaking the operator's language. Read it as a referee's reading, not a verdict: independent buys disinterest, not the last word. Captured 2026-07-27; dated sourcing on the research notes timeline.

Two disclosures this board earns. AA states it "supported Anthropic to evaluate Claude Opus 5 ahead of release" — its own harness and hardware, so not a vendor claim, but pre-release coordination on the row that came first. And the top of this board is a tie, not a ranking: AA's own stated confidence interval is ±1%, and the gap between the top two distinct models is 0.83 points. (AA ranks effort variants separately, so its literal second place is the same model again at a lower setting, 0.62 behind.) A page that discounts launch decks has to apply the same arithmetic to the referee.

Independent evals · Artificial Analysis
Agentic intelligence, priced per task
captured 2026-07-27 · Intelligence Index v4.1
captured today · fresh
Rank by
#ModelCost / taskIntel$/tasktok/s
1
DeepSeek V4 Pro (max)
44.3
$0.045
64
2
gpt-oss-120b (high)
23.8
$0.061
287
3
MiniMax-M3
44.4
$0.125
80
4
Nemotron 3 Ultra
37.8
$0.245
204
5
Muse Spark 1.1 (xhigh)
50.6
$0.261
128
6
GPT-5.6 Luna (max)
51.2
$0.277
188
7
GLM-5.2 (max)
51.1
$0.319
219
8
Grok 4.5 (high)
53.8
$0.350
56
9
Gemini 3.6 Flash
50.1
$0.501
235
10
Kimi K3
57.1
$0.723
33
11
GPT-5.6 Terra (max)
55.0
$0.825
136
12
Qwen3.7 Max
46.0
$1.03
202
13
Claude Sonnet 5 (max)
53.4
$1.52
74
14
GPT-5.6 Sol (max)
58.9
$1.54
77
15
Claude Opus 4.8 (max)
55.7
$1.80
55
16
Claude Opus 5 (max)
60.7
$2.03
53
17
Claude Fable 5
59.9
$2.75
71
Claude Opus 5 (max) tops the Index (60.7) at $2.03 per Index task — and is not the most expensive row on this board: Claude Fable 5 pays $2.75 for 0.8 fewer points. The frontier stopped being the priciest thing on the menu. Meanwhile DeepSeek V4 Pro (max) delivers 73% of the leader’s Index at 2.2% of its cost per task — a 45× price ratio. Capability is the vanity metric; cost-per-task is the one that shows up on the invoice. Sort by it.
How the Index is built — v4.1 (June 2026 — current)
  • Agents · 34%GDPval-AA v2 · 20% · 𝜏³-Banking · 14%
  • Coding · 24%Terminal-Bench v2.1 · 16% · SciCode · 8%
  • Scientific reasoning · 24%HLE · 12% · GPQA Diamond · 6% · CritPt · 6%
  • General · 18%AA-Omniscience · 12% (accuracy 8% + non-hallucination 4%) · AA-LCR · 6%
v4.1 retired IFBench (saturated), Terminal-Bench Hard, 𝜏²-Bench Telecom to chase agentic signal. It's a weighted composite — change the weights and you change the king. Independent buys disinterest, not infallibility: read it as a third reading that disagrees usefully with the crowd and the labs, not a tiebreaker that overrules them.
Precision. AA estimates a 95% confidence interval of less than ±1% on the Index. The gap between the top two distinct models — Opus 5 at max effort and Fable 5 — is 0.83 points, inside it. (AA ranks effort variants separately, so its literal #2 is Opus 5 at xhigh.) Treat the top of this board as a tie, not a ranking.
Disclosure. AA discloses it “supported Anthropic to evaluate Claude Opus 5 ahead of release.” Own harness, own hardware — but pre-release vendor coordination on the #1 row. Weigh it as such.
One column, one denominator. AA publishes two different numbers it calls "cost per task": the Intelligence Index one shown here, and a much larger AA-Briefcase one (Opus 5 at max: $17.79 a task, against Fable 5's $22.30). Secondary coverage quotes them interchangeably. This column is always the Index.
Source: Artificial Analysis — independent evals, run on their own harness. Hand-captured fair-use snapshot; figures verified against AA's public board on 2026-07-27. Methodology →Open artificialanalysis.ai →
What the labs claim

Launch-deck numbers — discounted on arrival

The arena above is crowd votes on chat prompts. The boards below are what vendors shipped on launch day for the question operators actually ask — agentic work — which the arena doesn't measure. Read every row as a claim, not a receipt: Berkeley RDI reward-hacked 8 of 8 major agent benchmarks, so launch numbers carry a 10–15 point discount until your own eval confirms them. Models you can't buy don't appear — capability ceilings are not tier-list entries. Dated sourcing for every figure lives on the research notes timeline.

Three of these four cards are vendor claims. One was run by a third party. Watch what happens to the numbers when the referee changes — that difference is the entire argument of this chapter, and it is now visible in a column rather than asserted in a sentence.

SWE-Bench Pro
Vendor claim
agentic coding — resolving real GitHub issues in real repos
claude-fable-5 80.0
claude-opus-5 79.2
claude-opus-4-8 69.2
gpt-5.6-sol 64.6
Anthropic · Claude Opus 5 System Card, Table 8.1.A · 2026-07-24
Max effort, mean of 5 trials — max is NOT the API default. Fable 5 still wins this row. Opus 5 appears on no independent SWE-Bench board. And Cursor's June study found 63% of Opus 4.8's passes retrieved a known fix rather than deriving one; sealing git history AND cutting internet access together dropped it 87.1% → 73.0% (of audited retrievals, web lookup was 57% and git-history mining only 9% — the internet is the bigger leak). Discount accordingly.
FrontierBench v0.1
Vendor claim
hard agentic terminal work — the Terminal-Bench successor
Anthropic's table Public board
claude-opus-5 43.3 43.5
gpt-5.6-sol 37.5 34.4
claude-fable-5 33.7 33.8
claude-opus-4-8 18.7 21.1
Anthropic System Card Table 8.1.A · 2026-07-24 — vs. frontierbench.ai public board, read 2026-07-27
The most instructive row on this page, and not for the reason you would expect: the independent board AGREES on Opus 5 (43.5 ±1.65, rank 1) and disagrees sharply about its rival — Anthropic re-ran GPT-5.6 Sol itself at 37.5, while the public board has Sol at 34.4 on OpenAI's own agent. Two caveats survive the agreement. Anthropic's separate §8.5 run scores best at xhigh (44.4), not the max its headline prints. And in that run Opus 4.8 silently substituted whenever a safety classifier refused — 4% of Opus 5's trials, 26% of Fable 5's — so both Anthropic scores are two-model blends while the OpenAI score is clean. Note also that the public board mixes harnesses: Opus 5's rank-1 row runs on Princeton's mini-SWE-agent, Fable 5's on Anthropic's own.
OSWorld 2.0
Vendor claim
computer use — long-horizon desktop automation
claude-opus-5 70.6
claude-fable-5 66.1
gpt-5.6-sol 62.6
claude-opus-4-8 55.7
Anthropic · Claude Opus 5 System Card, Table 8.1.A · 2026-07-24
This is PARTIAL/checkpoint credit, not task completion — the OSWorld 2.0 paper scores Opus 4.8 at 20.6% binary against 54.8% partial, so "70.6% on computer tasks" reads about 3× what it means. Opus 4.8 was the grader. The GPT-5.6 figure was lifted from OpenAI's own post, so the gap is cross-harness. Opus 5 is absent from the official board.
ARC-AGI-3
Independent
novel-environment reasoning, scored against human action efficiency
claude-opus-5 (high) 30.16%
anthropic fable-class ~20%
gpt-5.6-sol (max) 13.33%
ARC Prize Foundation · arcprize.org · 2026-07-24 — run by ARC Prize, not by Anthropic
The one launch-window number on this page that earns full credit, because a third party executed and scored it. Read the config honestly: ARC Prize evaluated HIGH effort only — "due to the short testing window" — so no max-effort figure exists. All three rows above are the same Public Demo set; Anthropic's own table mixes in a Semi-Private score for GPT-5.6 (7.8%), which inflates the gap from ~2.3× to ~4×. One unresolved discrepancy worth stating rather than smoothing: Anthropic's §8.14.2 describes these same results as semi-private, while ARC Prize's own results page files them under the Public Demo set. We follow ARC Prize, since ARC Prize ran it.
The board that measures the actual job

Arena's Agent board — and why it isn't a tab above

Arena runs an agent board too, over 1,242,857 sessions · 38 models. It is deliberately not one of the tabs above, because it does not use Elo — its metric is a Net Improvement Score, a percentage. Rendering 12.72 in a column that reads 1508 two tabs over would be the exact category error this chapter exists to argue against, so it gets its own shape. It is also the closest any crowd board comes to measuring what operators actually buy — and note who is missing from it.

Net improvement score board published 2026-07-21
1 Claude Fable 5 (High) 12.72% ± 2.00%
2 GPT 5.6 Sol (xHigh) 10.12% ± 1.69%
3 Claude Opus 4.8 (Thinking) 9.75% ± 1.39%
4 Kimi K3 9.71% ± 1.52%
5 Claude Sonnet 5 (High) 8.66% ± 1.89%
6 GPT 5.5 (xHigh) 8.41% ± 0.87%
7 Claude Opus 4.7 (Thinking) 7.94% ± 1.24%
8 Claude Opus 4.7 7.67% ± 1.25%
9 GPT 5.5 (High) 7.61% ± 0.81%
10 GLM 5.2 (Max) 6.50% ± 1.00%
Opus 5 is not on this board at all — three days old at capture. Every interval here overlaps its neighbour's, so read it as three or four bands, not ten ranks. The band that matters: an open-weights model (Kimi K3) sits inside the top four, and the #3 slot belongs to a model one generation old.
What an operator actually uses

Vlad's tier list

The arena measures capability head-to-head on neutral prompts. This list measures what runs Vlad's portfolio on a Tuesday at 11am. They disagree more than you'd expect — Claude Code and Cowork don't appear on the arena at all, Perplexity loses head-to-head but is S-tier here because it answers the actual question 90% of the time. Drag tools, build your own, share the URL.

Build your own tier list
S
Run my life — remove this and three things break by Wednesday
A
Open every day
B
Useful for one job each
C
I see why people use these but I don't
D
Exists, fine, not for me
F
Actively bad / don't
·
Unranked pool — drag into a tier
Drag tools between tiers. State saves locally. Share opens copy / Tweet / LinkedIn / device share — the URL encodes every placement, so whoever opens it sees exactly your tiers.
The argument

Benchmarks reward the model that beats other models. Operators reward the tool that doesn't break the workflow. The most useful tool you own is rarely the most capable one — it's the one with the lowest activation energy on a Tuesday morning when you have nineteen other problems. That's why Claude Code is S-tier in this list and not even ranked on LMArena. Different question. Different answer.

Stay close

The next edition lands when this list says it does.

No course. No paywall. Operator playbooks weekly. 10K+ subscribers.