Chapter 24

The tier list. Without mercy.

Public leaderboards rank performance under specific conditions. Operators rank usefulness. Start with the current model notes, then compare crowd votes, independent task economics, archived lab claims, agent results and Vlad's personal tiers. These are different questions, not five votes for one universal winner.

Captured evidence, not a live feed

Arena captured ; Artificial Analysis captured (Intelligence Index v4.2). Each Arena board carries its own vote cutoff. Capture dates describe when we read a source, not when every result was measured.

Keep effort settings, harnesses and fallback conditions attached to every score. New Index versions are not numerically comparable with earlier versions, and an unavailable source does not justify carrying an old score forward as current. The lab comparison archive, the 2026-09-28 model tiers and the July 27 tools tiers are each dated separately.

Model releases checked 2026-09-05

Astra and Fable 5.1

Release notes, not personal tiers; Vlad's model tiers are in the operator section. The workflow candidates below are editorial hypotheses based on documented capabilities. Dive's matched-workload evaluation is pending.

GPT-6 Astra

Released

gpt-6-astra

Staged rollout. Availability differs across ChatGPT, Work, Codex, and the API; check your plan and workspace.

API context / maximum output
1,050,000-token context; 128,000-token maximum output.
API input / output
$10 / $50 per million tokens. Standard API rate up to 272K input tokens. Above that threshold, the whole request costs $20 input / $75 output per million tokens. Cache and tool charges are separate.
Work to evaluate
Scoped repository changes, source-grounded research, and document creation with explicit acceptance tests.
Comparison boundary
Keep effort and harness versions in every result. Launch scores do not isolate model changes from changes to the agent system.

Claude Fable 5.1

Released

claude-fable-5-1

Generally available. Included allowance versus usage credits depends on the Claude plan; the Fable 5 launch promotion does not apply.

API context / maximum output
1,000,000-token context; 128,000-token maximum output.
API input / output
$10 / $50 per million tokens. Standard Claude API rate across the 1M context. Cache reads cost $0.25 per million tokens; cache writes and tool charges are separate. Provider rates may differ.
Work to evaluate
Difficult debugging, multi-stage engineering, and research-to-document workflows with review checkpoints.
Comparison boundary
Record fallback and the model that actually served the task. Restricted-access Mythos 5.1 has different domain safeguards; its benchmark scores are not interchangeable with Fable 5.1 scores.

Put the reference into a trial: GPT-6 Astra field guide and Claude Fable 5.1 field guide. These are source-backed guides with proposed exercises, not completed comparative tests. The Claude Opus 5.5 and Claude Sonnet 5.5 field guides add one small same-task measurement, two runs per setting.

What the public says

Arena public leaderboard

Crowdsourced head-to-head model votes from arena.ai, formerly LMArena. Each board measures preference within its own task category and voting pool. A top row is not a guarantee of better results in your repository or workflow.

Read the board's date, vote coverage and confidence intervals before treating a rank gap as a meaningful difference. Model suffixes matter: effort levels, harnesses and preview variants are distinct configurations, not interchangeable names. An absent model is unmeasured here, not proven worse.

Arena · Text
The headline board — head-to-head chat votes.
captured 2026-09-05
votes through 2026-09-02 · 8.00M votes · 400 models
1
Anthropic
claude-fable-5
1507
2
Anthropic
claude-opus-4-6-high
1505
3
Anthropic
claude-fable-5.1-max
1504
4
Anthropic
claude-opus-4-7-high
1502
5
Meta
muse-spark-1.2 (xHigh)
1499
6
Anthropic
claude-opus-4-6
1498
7
Anthropic
claude-opus-4-7
1494
8
Google
gemini-3.8-flash-high
1494
9
Anthropic
claude-opus-5-high
1493
10
Meta
muse-spark-1.1
1492
11
Google
gemini-3.7-flash-high
1491
12
Moonshot
kimi-k3-max
1489
Anthropic holds 7 of the top 10 here. The Elo spread across this top 10 is 15 points. Elo gaps are specific to the Text board; the spread alone does not establish a statistically resolved ranking or predict performance on your tasks.
Read the board, not the rank. Style-controlled overall board. Fable 5.1 enters at #3 as claude-fable-5.1-max: 1504 on 2,906 votes, with a 1493–1515 interval that overlaps Fable 5 at #1 (1502–1512). This is a max-effort result, not a generic Fable 5.1 score. Gemini 3.8 Flash and 3.7 Flash carry Arena's Preliminary flag. No OpenAI model is in the top 12; GPT-6 Astra is catalog-listed but has no published row on this board.
Captured 2026-09-05 from Arena's published board. Ten of eleven Elo boards were cross-checked against the 2026-09-05 community mirror capture; Image-to-Code is single-sourced. Arena remains the primary source.Open Text on arena.ai →
What an independent referee measures

Artificial Analysis — the independent referee

Artificial Analysis provides an independent evaluation and reported task costs, separate from crowd preferences and vendor launch tables. This economics subset includes models with both an Intelligence Index score and a cost per Index task. Captured 2026-09-05, using Intelligence Index v4.2; the panel carries the methodology and source disclosures.

Compare within this Index version, not against older captures. A changed benchmark mix or weighting changes what the score means. Keep the exact effort and fallback labels when comparing rows, distinguish unavailable metrics from zero, and do not interpret small score gaps without uncertainty. A cheaper Index task is not automatically a cheaper successful workflow.

Independent evals · Artificial Analysis
Intelligence, priced per task
captured 2026-09-05 · Intelligence Index v4.2
Dated public snapshot
Rank by
#ModelCost / taskIntel$/tasktok/s
1
GPT-5.6 Luna (max)
43.4
$0.097
133
2
Gemini 3.8 Flash (high)
47.1
$0.738
—
3
Muse Spark 1.3 (max)
53.0
$0.959
190
4
GPT-5.6 Sol (max)
51.3
$1.25
85
5
GPT-6 Astra (max)
54.7
$2.57
63
6
Claude Opus 5 (max)
54.1
$4.21
58
7
Claude Fable 5.1 (max with fallback)
56.8
$6.12
67
Claude Fable 5.1 (max with fallback) has the highest Index score in this selection (56.8) at $6.12 per Index task. GPT-5.6 Luna (max) is the lowest-cost selected row: 43.4 points at $0.097 per task, a 63× cost difference. These are weighted benchmark costs, not a quote for your workload.

Seven selected model settings with a verified Index score and task cost, not the full leaderboard. Effort and fallback labels are preserved. AA mentions Astra xhigh in its page summary but supplies no numeric xhigh row in the captured JSON-LD; max is not a substitute. Missing values are unverified, not zero.

How the Index is built — v4.2 (September 2026; announced September 4)
  • Agents · 30% — AA-Briefcase · 15% · GDPval-AA v2 · 10% · 𝜏³-Banking · 5%
  • Coding · 20% — Terminal-Bench v2.1 · 10% · SciCode · 10%
  • Scientific reasoning · 20% — HLE · 10% · CritPt · 10%
  • General · 30% — AA-Omniscience · 15% (accuracy 10% + non-hallucination 5%) · GDP.pdf · 10% · AA-LCR v1.1 · 5%
Ten evaluations: v4.2 adds AA-Briefcase and GDP.pdf, removes GPQA Diamond from the Index, upgrades AA-LCR to v1.1 and regrades SciCode v1.0.1. Weights and grading changed, so scores and per-task costs are not a like-for-like trend against July or August captures.
Precision. AA estimates a 95% confidence interval below ±1% for the Index; individual benchmarks may be wider. That aggregate estimate is not a published pairwise significance test. Small score differences do not establish a universal winner.
Disclosure. These are AA evaluations, not vendor claims or our own tests. Fable 5.1 is measured with fallback, as served. The public snapshot does not establish whether any launch involved pre-release coordination; the July Opus 5 disclosure is not evidence about the current leader.
Agentic evidence. The former Agentic Index URL returned HTTP 404 at this capture. No comparable current standalone Agentic scores were verified in the public JSON-LD; the old column is omitted, not carried into v4.2 or reconstructed from individual benchmarks.
Speed. Speed is the /models chart output rate, not end-to-end task time. Astra max differs between that chart and its detail-page summary at capture; this panel keeps the chart value. Gemini 3.8 Flash has no speed in the captured chart, so it remains unscored here.
One denominator. Cost is the weighted USD average per Intelligence Index task, including input, answer, reasoning and cache tokens. It is not standalone AA-Briefcase cost, a whole-Index run, or token-list pricing.
Source: Artificial Analysis · Artificial Analysis (2025). LLM benchmarks dataset. Limited public-data selection captured 2026-09-05; no keyed API. Methodology → · Source termsOpen artificialanalysis.ai →
Dated comparison archive

Lab claims, with their test conditions

Historical comparison captured 2026-07-27, not the latest model ranking. These cards preserve the Opus 5 launch-window claims and independent comparisons available then. Their scores, absences and commentary belong to that archive; Astra and Fable 5.1 are covered in the current model notes.

Vendor-run and independently run results carry different provenance, not an automatic point adjustment. Check the dataset split, grader, tool access, effort setting and fallback model before trusting a comparison. Where those conditions differ, the gap is not a clean model-to-model result. Validate on your own workload; dated sources remain on the research notes timeline.

Comparison archive · 2026-07-27
SWE-Bench Pro
Vendor claim
agentic coding — resolving real GitHub issues in real repos
claude-fable-5 80.0
claude-opus-5 79.2
claude-opus-4-8 69.2
gpt-5.6-sol 64.6
Anthropic · Claude Opus 5 System Card, Table 8.1.A · 2026-07-24
Max effort, mean of 5 trials — max is NOT the API default. Fable 5 still wins this row. Opus 5 appears on no independent SWE-Bench board. And Cursor's June study found 63% of Opus 4.8's passes retrieved a known fix rather than deriving one; sealing git history AND cutting internet access together dropped it 87.1% → 73.0% (of audited retrievals, web lookup was 57% and git-history mining only 9% — the internet is the bigger leak). Discount accordingly.
Comparison archive · 2026-07-27
FrontierBench v0.1
Vendor claim
hard agentic terminal work — the Terminal-Bench successor
Anthropic's table Public board
claude-opus-5 43.3 43.5
gpt-5.6-sol 37.5 34.4
claude-fable-5 33.7 33.8
claude-opus-4-8 18.7 21.1
Anthropic System Card Table 8.1.A · 2026-07-24 — vs. frontierbench.ai public board, read 2026-07-27
The most instructive row on this page, and not for the reason you would expect: the independent board AGREES on Opus 5 (43.5 ±1.65, rank 1) and disagrees sharply about its rival — Anthropic re-ran GPT-5.6 Sol itself at 37.5, while the public board has Sol at 34.4 on OpenAI's own agent. Two caveats survive the agreement. Anthropic's separate §8.5 run scores best at xhigh (44.4), not the max its headline prints. And in that run Opus 4.8 silently substituted whenever a safety classifier refused — 4% of Opus 5's trials, 26% of Fable 5's — so both Anthropic scores are two-model blends while the OpenAI score is clean. Note also that the public board mixes harnesses: Opus 5's rank-1 row runs on Princeton's mini-SWE-agent, Fable 5's on Anthropic's own.
Comparison archive · 2026-07-27
OSWorld 2.0
Vendor claim
computer use — long-horizon desktop automation
claude-opus-5 70.6
claude-fable-5 66.1
gpt-5.6-sol 62.6
claude-opus-4-8 55.7
Anthropic · Claude Opus 5 System Card, Table 8.1.A · 2026-07-24
This is PARTIAL/checkpoint credit, not task completion — the OSWorld 2.0 paper scores Opus 4.8 at 20.6% binary against 54.8% partial, so "70.6% on computer tasks" reads about 3× what it means. Opus 4.8 was the grader. The GPT-5.6 figure was lifted from OpenAI's own post, so the gap is cross-harness. Opus 5 is absent from the official board.
Comparison archive · 2026-07-27
ARC-AGI-3
Independent
novel-environment reasoning, scored against human action efficiency
claude-opus-5 (high) 30.16%
anthropic fable-class ~20%
gpt-5.6-sol (max) 13.33%
ARC Prize Foundation · arcprize.org · 2026-07-24 — run by ARC Prize, not by Anthropic
The one launch-window number on this page that earns full credit, because a third party executed and scored it. Read the config honestly: ARC Prize evaluated HIGH effort only — "due to the short testing window" — so no max-effort figure exists. All three rows above are the same Public Demo set; Anthropic's own table mixes in a Semi-Private score for GPT-5.6 (7.8%), which inflates the gap from ~2.3× to ~4×. One unresolved discrepancy worth stating rather than smoothing: Anthropic's §8.14.2 describes these same results as semi-private, while ARC Prize's own results page files them under the Public Demo set. We follow ARC Prize, since ARC Prize ran it.
A separate measure of agent work

Arena's Agent board — and why it isn't a tab above

Arena's agent results use Net Improvement Score, not the Elo scale of the crowd boards. The published snapshot covers 2,188,416 sessions · 58 models, captured 2026-09-05. These scores describe agent performance under that board's conditions, not a direct comparison with the Intelligence Index or a substitute for testing your own workflow.

Net Improvement Score board published 2026-09-01
1 Claude Opus 5 (High) 13.74% ± 1.80%
2 Claude Opus 5 (Max) 11.69% ± 2.01%
3 Claude Fable 5 (High) 10.61% ± 1.53%
4 GPT 5.6 Sol (xHigh) 9.49% ± 1.51%
5 Claude Opus 4.8 (High) 9.22% ± 1.53%
6 Kimi K3 (Max) 8.71% ± 0.66%
7 GPT 5.5 (xHigh) 7.53% ± 1.08%
8 Claude Sonnet 5 (High) 7.51% ± 2.11%
9 Claude Opus 4.7 (High) 6.49% ± 1.42%
10 GLM 5.2 (Max) 6.23% ± 0.77%

Overall Agent board, not its Code subcategory. Percentages are 100 × avgScore.value with 100 × avgScore.ci, not Elo or a task-success rate. Opus 5 high leads at 13.74% ± 1.80%, with an interval overlapping both Opus 5 max and Fable 5 high. Neither Fable 5.1 nor GPT-6 Astra appears among the 58 published rows; shared model-catalog entries are not ranked evidence.

Read the displayed intervals alongside rank and retain each row's effort setting. Net improvement is relative to the model pool: changes in that pool can move scores without a change in the model itself. Neither a raw-score delta nor a rank change across captures establishes a capability gain.

What an operator actually uses

Vlad's tier list

Model tiers updated ; tools tiers from 2026-07-27. Both are Vlad's workflow preferences, not benchmark rankings. A placement here is not a matched-workload test; Dive's evaluation of Astra and Fable 5.1 is still pending.

Vlad's model tiers Updated
  1. Tier SSS
    • Opus 5.5
  2. Tier SS
    • GPT-6 Astra
  3. Tier S
    • Fable 5.1
    • GPT-6 Sol
  4. Tier A
    • Opus 5
    • Fable 5
    • GPT-5.6 SOL
    • Kimi K3
    • Grok 4.6
    • Qwen3.8 Max 0902
    • GLM-5.3
    • GPT-6 Luna
    • Muse Spark 1.3
    • DeepSeek V4.1 Flash
    • Hy-4 Preview
    • Grok 4.7
  5. Tier B
    • GPT-5.6 Terra
    • Qwen3.8-Flash-Next
    • GLM-5.3 Flash
    • Sonnet 5
    • K2 Horizon 375B A23B
  6. Tier C
    • GPT-5.6 Luna
    • Grok 4.5
    • Qwen3.8 27B
    • Muse Spark 1.2
    • mimo V2.5 Pro
    • Hy-3
    • MiniMax-M3
  7. Tier D
    • Inkling
    • Nemotron 3 Ultra
    • Muse Glimmer
  8. Tier E
    • Haiku 4.5
    • Mistral Medium 3.5
    • Nemotron 3.5 Lightning
  9. Tier Google
    • Gemini 3.8 Flash
    • Gemini 3.7 Flash
    • Gemini 3.6 Flash
    • Gemini 3.5 Flash
Personal placements, not a benchmark ranking. Dots use the Arena board's lab colours; labs outside that key have none.

Tools and surfaces

Tools snapshot: 2026-07-27. Drag them into your own order, then share the link.

Build your own tier list
S
Run my life — remove this and three things break by Wednesday
A
Open every day
B
Useful for one job each
C
I see why people use these but I don't
D
Exists, fine, not for me
F
Actively bad / don't
·
Unranked pool — drag into a tier
Drag tools between tiers. State saves locally. Share opens copy / Tweet / LinkedIn / device share — the URL encodes every placement, so whoever opens it sees exactly your tiers.

Reading the evidence

What does an independent benchmark establish?
Artificial Analysis provides an evaluation separate from vendor launch claims and crowd preference votes. Independence does not remove uncertainty: check the harness, dataset version, model configuration and any disclosed vendor coordination before applying the result to your work.
What does cost per task mean here?
It is the reported US-dollar cost to run an Artificial Analysis Intelligence Index task, not a token list price or the cost of a successful production job. Retries, tool calls, latency and human review can change your workflow economics.
Does a leaderboard rank decide the operator tier?
No. Rank describes one board and tested configuration. Effort settings, tools, harnesses and safety fallbacks can change what was actually evaluated. Match those conditions, then validate quality, cost and reliability on your own tasks. Vlad's model tiers were updated on 2026-09-28 and the tools defaults remain the July 27, 2026 snapshot; both are personal placements, not matched-workload results.
Which captured models lead on score and task cost?
In our 2026-09-05 capture of Artificial Analysis's Intelligence Index v4.2, the highest listed score in this economics subset is 56.8 (Claude Fable 5.1 (max with fallback)); the lowest listed cost is $0.097 per Intelligence Index task (GPT-5.6 Luna (max)). Only models with both a score and a task cost are included. These are captured configurations, not a universal winner or a quote for your workflow.

Score and cost source: Artificial Analysis, captured 2026-09-05. Compare the exact configurations in the independent panel.

The argument

Benchmarks reward the model that beats other models. Operators reward the tool that doesn't break the workflow. The most useful tool you own is rarely the most capable one — it's the one with the lowest activation energy on a Tuesday morning when you have nineteen other problems. That's why Claude Code is S-tier in the tools list and not even ranked on LMArena. Different question. Different answer.

Stay close

The next edition lands when this list says it does.

No course. No paywall. Operator playbooks weekly. 10K+ subscribers.