Public leaderboards rank performance under specific conditions. Operators rank usefulness. Start with the current model notes, then compare crowd votes, independent task economics, archived lab claims, agent results and Vlad's personal tiers. These are different questions, not five votes for one universal winner.
Captured evidence, not a live feed
Arena captured ; Artificial Analysis captured (Intelligence Index v4.2). Each Arena board carries its own vote cutoff. Capture dates describe when we read a source, not when every result was measured.
Keep effort settings, harnesses and fallback conditions attached to every score. New Index versions are not numerically comparable with earlier versions, and an unavailable source does not justify carrying an old score forward as current. The lab comparison archive, the 2026-09-28 model tiers and the July 27 tools tiers are each dated separately.
Model releases checked 2026-09-05
Astra and Fable 5.1
Release notes, not personal tiers; Vlad's model tiers are in the operator section. The workflow candidates below are editorial hypotheses based on documented capabilities. Dive's matched-workload evaluation is pending.
GPT-6 Astra
Released
gpt-6-astra
Staged rollout. Availability differs across ChatGPT, Work, Codex, and the API; check your plan and workspace.
API context / maximum output
1,050,000-token context; 128,000-token maximum output.
API input / output
$10 / $50 per million tokens. Standard API rate up to 272K input tokens. Above that threshold, the whole request costs $20 input / $75 output per million tokens. Cache and tool charges are separate.
Work to evaluate
Scoped repository changes, source-grounded research, and document creation with explicit acceptance tests.
Comparison boundary
Keep effort and harness versions in every result. Launch scores do not isolate model changes from changes to the agent system.
Generally available. Included allowance versus usage credits depends on the Claude plan; the Fable 5 launch promotion does not apply.
API context / maximum output
1,000,000-token context; 128,000-token maximum output.
API input / output
$10 / $50 per million tokens. Standard Claude API rate across the 1M context. Cache reads cost $0.25 per million tokens; cache writes and tool charges are separate. Provider rates may differ.
Work to evaluate
Difficult debugging, multi-stage engineering, and research-to-document workflows with review checkpoints.
Comparison boundary
Record fallback and the model that actually served the task. Restricted-access Mythos 5.1 has different domain safeguards; its benchmark scores are not interchangeable with Fable 5.1 scores.
Crowdsourced head-to-head model votes from arena.ai, formerly LMArena. Each board measures preference within its own task category and voting pool. A top row is not a guarantee of better results in your repository or workflow.
Read the board's date, vote coverage and confidence intervals before treating a rank gap as a meaningful difference. Model suffixes matter: effort levels, harnesses and preview variants are distinct configurations, not interchangeable names. An absent model is unmeasured here, not proven worse.
Arena · Text
The headline board — head-to-head chat votes.
captured 2026-09-05
votes through 2026-09-02 · 8.00M votes · 400 models
1
Anthropic
claude-fable-5
1507
2
Anthropic
claude-opus-4-6-high
1505
3
Anthropic
claude-fable-5.1-max
1504
4
Anthropic
claude-opus-4-7-high
1502
5
Meta
muse-spark-1.2 (xHigh)
1499
6
Anthropic
claude-opus-4-6
1498
7
Anthropic
claude-opus-4-7
1494
8
Google
gemini-3.8-flash-high
1494
9
Anthropic
claude-opus-5-high
1493
10
Meta
muse-spark-1.1
1492
11
Google
gemini-3.7-flash-high
1491
12
Moonshot
kimi-k3-max
1489
Anthropic holds 7 of the top 10 here. The Elo spread across this top 10 is 15 points. Elo gaps are specific to the Text board; the spread alone does not establish a statistically resolved ranking or predict performance on your tasks.
Read the board, not the rank. Style-controlled overall board. Fable 5.1 enters at #3 as claude-fable-5.1-max: 1504 on 2,906 votes, with a 1493–1515 interval that overlaps Fable 5 at #1 (1502–1512). This is a max-effort result, not a generic Fable 5.1 score. Gemini 3.8 Flash and 3.7 Flash carry Arena's Preliminary flag. No OpenAI model is in the top 12; GPT-6 Astra is catalog-listed but has no published row on this board.
Captured 2026-09-05 from Arena's published board. Ten of eleven Elo boards were cross-checked against the 2026-09-05 community mirror capture; Image-to-Code is single-sourced. Arena remains the primary source.Open Text on arena.ai →
What an independent referee measures
Artificial Analysis — the independent referee
Artificial Analysis provides an independent evaluation and reported task costs, separate from crowd preferences and vendor launch tables. This economics subset includes models with both an Intelligence Index score and a cost per Index task. Captured 2026-09-05, using Intelligence Index v4.2; the panel carries the methodology and source disclosures.
Compare within this Index version, not against older captures. A changed benchmark mix or weighting changes what the score means. Keep the exact effort and fallback labels when comparing rows, distinguish unavailable metrics from zero, and do not interpret small score gaps without uncertainty. A cheaper Index task is not automatically a cheaper successful workflow.
Independent evals · Artificial Analysis
Intelligence, priced per task
captured 2026-09-05 · Intelligence Index v4.2
Dated public snapshot
Rank by
#ModelCost / taskIntel$/tasktok/s
1
GPT-5.6 Luna (max)
43.4
$0.097
133
2
Gemini 3.8 Flash (high)
47.1
$0.738
—
3
Muse Spark 1.3 (max)
53.0
$0.959
190
4
GPT-5.6 Sol (max)
51.3
$1.25
85
5
GPT-6 Astra (max)
54.7
$2.57
63
6
Claude Opus 5 (max)
54.1
$4.21
58
7
Claude Fable 5.1 (max with fallback)
56.8
$6.12
67
Claude Fable 5.1 (max with fallback) has the highest Index score in this selection (56.8) at $6.12 per Index task. GPT-5.6 Luna (max) is the lowest-cost selected row: 43.4 points at $0.097 per task, a 63× cost difference. These are weighted benchmark costs, not a quote for your workload.
Seven selected model settings with a verified Index score and task cost, not the full leaderboard. Effort and fallback labels are preserved. AA mentions Astra xhigh in its page summary but supplies no numeric xhigh row in the captured JSON-LD; max is not a substitute. Missing values are unverified, not zero.
How the Index is built — v4.2 (September 2026; announced September 4)
Ten evaluations: v4.2 adds AA-Briefcase and GDP.pdf, removes GPQA Diamond from the Index, upgrades AA-LCR to v1.1 and regrades SciCode v1.0.1. Weights and grading changed, so scores and per-task costs are not a like-for-like trend against July or August captures.
Precision. AA estimates a 95% confidence interval below ±1% for the Index; individual benchmarks may be wider. That aggregate estimate is not a published pairwise significance test. Small score differences do not establish a universal winner.
Disclosure. These are AA evaluations, not vendor claims or our own tests. Fable 5.1 is measured with fallback, as served. The public snapshot does not establish whether any launch involved pre-release coordination; the July Opus 5 disclosure is not evidence about the current leader.
Agentic evidence. The former Agentic Index URL returned HTTP 404 at this capture. No comparable current standalone Agentic scores were verified in the public JSON-LD; the old column is omitted, not carried into v4.2 or reconstructed from individual benchmarks.
Speed. Speed is the /models chart output rate, not end-to-end task time. Astra max differs between that chart and its detail-page summary at capture; this panel keeps the chart value. Gemini 3.8 Flash has no speed in the captured chart, so it remains unscored here.
One denominator. Cost is the weighted USD average per Intelligence Index task, including input, answer, reasoning and cache tokens. It is not standalone AA-Briefcase cost, a whole-Index run, or token-list pricing.
Historical comparison captured 2026-07-27, not the latest model ranking. These cards preserve the Opus 5 launch-window claims and independent comparisons available then. Their scores, absences and commentary belong to that archive; Astra and Fable 5.1 are covered in the current model notes.
Vendor-run and independently run results carry different provenance, not an automatic point adjustment. Check the dataset split, grader, tool access, effort setting and fallback model before trusting a comparison. Where those conditions differ, the gap is not a clean model-to-model result. Validate on your own workload; dated sources remain on the research notes timeline.
Comparison archive · 2026-07-27
SWE-Bench Pro
Vendor claim
agentic coding — resolving real GitHub issues in real repos
claude-fable-580.0
claude-opus-579.2
claude-opus-4-869.2
gpt-5.6-sol64.6
Anthropic · Claude Opus 5 System Card, Table 8.1.A · 2026-07-24 Max effort, mean of 5 trials — max is NOT the API default. Fable 5 still wins this row. Opus 5 appears on no independent SWE-Bench board. And Cursor's June study found 63% of Opus 4.8's passes retrieved a known fix rather than deriving one; sealing git history AND cutting internet access together dropped it 87.1% → 73.0% (of audited retrievals, web lookup was 57% and git-history mining only 9% — the internet is the bigger leak). Discount accordingly.
Comparison archive · 2026-07-27
FrontierBench v0.1
Vendor claim
hard agentic terminal work — the Terminal-Bench successor
Anthropic's tablePublic board
claude-opus-543.343.5
gpt-5.6-sol37.534.4
claude-fable-533.733.8
claude-opus-4-818.721.1
Anthropic System Card Table 8.1.A · 2026-07-24 — vs. frontierbench.ai public board, read 2026-07-27 The most instructive row on this page, and not for the reason you would expect: the independent board AGREES on Opus 5 (43.5 ±1.65, rank 1) and disagrees sharply about its rival — Anthropic re-ran GPT-5.6 Sol itself at 37.5, while the public board has Sol at 34.4 on OpenAI's own agent. Two caveats survive the agreement. Anthropic's separate §8.5 run scores best at xhigh (44.4), not the max its headline prints. And in that run Opus 4.8 silently substituted whenever a safety classifier refused — 4% of Opus 5's trials, 26% of Fable 5's — so both Anthropic scores are two-model blends while the OpenAI score is clean. Note also that the public board mixes harnesses: Opus 5's rank-1 row runs on Princeton's mini-SWE-agent, Fable 5's on Anthropic's own.
Comparison archive · 2026-07-27
OSWorld 2.0
Vendor claim
computer use — long-horizon desktop automation
claude-opus-570.6
claude-fable-566.1
gpt-5.6-sol62.6
claude-opus-4-855.7
Anthropic · Claude Opus 5 System Card, Table 8.1.A · 2026-07-24 This is PARTIAL/checkpoint credit, not task completion — the OSWorld 2.0 paper scores Opus 4.8 at 20.6% binary against 54.8% partial, so "70.6% on computer tasks" reads about 3× what it means. Opus 4.8 was the grader. The GPT-5.6 figure was lifted from OpenAI's own post, so the gap is cross-harness. Opus 5 is absent from the official board.
Comparison archive · 2026-07-27
ARC-AGI-3
Independent
novel-environment reasoning, scored against human action efficiency
claude-opus-5 (high)30.16%
anthropic fable-class~20%
gpt-5.6-sol (max)13.33%
ARC Prize Foundation · arcprize.org · 2026-07-24 — run by ARC Prize, not by Anthropic The one launch-window number on this page that earns full credit, because a third party executed and scored it. Read the config honestly: ARC Prize evaluated HIGH effort only — "due to the short testing window" — so no max-effort figure exists. All three rows above are the same Public Demo set; Anthropic's own table mixes in a Semi-Private score for GPT-5.6 (7.8%), which inflates the gap from ~2.3× to ~4×. One unresolved discrepancy worth stating rather than smoothing: Anthropic's §8.14.2 describes these same results as semi-private, while ARC Prize's own results page files them under the Public Demo set. We follow ARC Prize, since ARC Prize ran it.
A separate measure of agent work
Arena's Agent board — and why it isn't a tab above
Arena's agent results use Net Improvement Score, not the Elo scale of the crowd boards. The published snapshot covers 2,188,416 sessions · 58 models, captured 2026-09-05. These scores describe agent performance under that board's conditions, not a direct comparison with the Intelligence Index or a substitute for testing your own workflow.
Net Improvement Scoreboard published 2026-09-01
1Claude Opus 5 (High)13.74% ± 1.80%
2Claude Opus 5 (Max)11.69% ± 2.01%
3Claude Fable 5 (High)10.61% ± 1.53%
4GPT 5.6 Sol (xHigh)9.49% ± 1.51%
5Claude Opus 4.8 (High)9.22% ± 1.53%
6Kimi K3 (Max)8.71% ± 0.66%
7GPT 5.5 (xHigh)7.53% ± 1.08%
8Claude Sonnet 5 (High)7.51% ± 2.11%
9Claude Opus 4.7 (High)6.49% ± 1.42%
10GLM 5.2 (Max)6.23% ± 0.77%
Overall Agent board, not its Code subcategory. Percentages are 100 × avgScore.value with 100 × avgScore.ci, not Elo or a task-success rate. Opus 5 high leads at 13.74% ± 1.80%, with an interval overlapping both Opus 5 max and Fable 5 high. Neither Fable 5.1 nor GPT-6 Astra appears among the 58 published rows; shared model-catalog entries are not ranked evidence.
Read the displayed intervals alongside rank and retain each row's effort setting. Net improvement is relative to the model pool: changes in that pool can move scores without a change in the model itself. Neither a raw-score delta nor a rank change across captures establishes a capability gain.
What an operator actually uses
Vlad's tier list
Model tiers updated ; tools tiers from 2026-07-27. Both are Vlad's workflow preferences, not benchmark rankings. A placement here is not a matched-workload test; Dive's evaluation of Astra and Fable 5.1 is still pending.
Vlad's model tiersUpdated
Tier SSS
Opus 5.5
Tier SS
GPT-6 Astra
Tier S
Fable 5.1
GPT-6 Sol
Tier A
Opus 5
Fable 5
GPT-5.6 SOL
Kimi K3
Grok 4.6
Qwen3.8 Max 0902
GLM-5.3
GPT-6 Luna
Muse Spark 1.3
DeepSeek V4.1 Flash
Hy-4 Preview
Grok 4.7
Tier B
GPT-5.6 Terra
Qwen3.8-Flash-Next
GLM-5.3 Flash
Sonnet 5
K2 Horizon 375B A23B
Tier C
GPT-5.6 Luna
Grok 4.5
Qwen3.8 27B
Muse Spark 1.2
mimo V2.5 Pro
Hy-3
MiniMax-M3
Tier D
Inkling
Nemotron 3 Ultra
Muse Glimmer
Tier E
Haiku 4.5
Mistral Medium 3.5
Nemotron 3.5 Lightning
Tier Google
Gemini 3.8 Flash
Gemini 3.7 Flash
Gemini 3.6 Flash
Gemini 3.5 Flash
Personal placements, not a benchmark ranking. Dots use the Arena board's lab colours; labs outside that key have none.
Tools and surfaces
Tools snapshot: 2026-07-27. Drag them into your own order, then share the link.
Build your own tier list
S
Run my life — remove this and three things break by Wednesday
A
Open every day
B
Useful for one job each
C
I see why people use these but I don't
D
Exists, fine, not for me
F
Actively bad / don't
·
Unranked pool — drag into a tier
Drag tools between tiers. State saves locally. Share opens copy / Tweet / LinkedIn / device share — the URL encodes every placement, so whoever opens it sees exactly your tiers.
Reading the evidence
What does an independent benchmark establish?
Artificial Analysis provides an evaluation separate from vendor launch claims and crowd preference votes. Independence does not remove uncertainty: check the harness, dataset version, model configuration and any disclosed vendor coordination before applying the result to your work.
What does cost per task mean here?
It is the reported US-dollar cost to run an Artificial Analysis Intelligence Index task, not a token list price or the cost of a successful production job. Retries, tool calls, latency and human review can change your workflow economics.
Does a leaderboard rank decide the operator tier?
No. Rank describes one board and tested configuration. Effort settings, tools, harnesses and safety fallbacks can change what was actually evaluated. Match those conditions, then validate quality, cost and reliability on your own tasks. Vlad's model tiers were updated on 2026-09-28 and the tools defaults remain the July 27, 2026 snapshot; both are personal placements, not matched-workload results.
Which captured models lead on score and task cost?
In our 2026-09-05 capture of Artificial Analysis's Intelligence Index v4.2, the highest listed score in this economics subset is 56.8 (Claude Fable 5.1 (max with fallback)); the lowest listed cost is $0.097 per Intelligence Index task (GPT-5.6 Luna (max)). Only models with both a score and a task cost are included. These are captured configurations, not a universal winner or a quote for your workflow.
Score and cost source: Artificial Analysis, captured 2026-09-05. Compare the exact configurations in the independent panel.
The argument
Benchmarks reward the model that beats other models. Operators reward the tool that doesn't break the workflow. The most useful tool you own is rarely the most capable one — it's the one with the lowest activation energy on a Tuesday morning when you have nineteen other problems. That's why Claude Code is S-tier in the tools list and not even ranked on LMArena. Different question. Different answer.
Stay close
The next edition lands when this list says it does.
No course. No paywall. Operator playbooks weekly. 10K+ subscribers.