Edition·2026-09-20 · Sun
Fed from·Epoch AI·HF Open LLM·OpenAlex·OpenAI evals
The benchmark record for AI models.
Every reported score is tagged by how it was checked — independently reproduced, vendor self-reported, contradicted, or unverified — and every benchmark is graded on the four signals that decide whether a number still means something. We grade the evaluations, not just the models.
Latest measurement·2026-09-03·34 of 62 benchmarks graded A
How trustworthy the evaluations are
Benchmark integrity
Trust holds the top-left; the bottom-right is where scores max out and belief erodes. The best-graded benchmarks:
A | Chess Puzzles | consistent harness |
A | FrontierMath-Tiers-1-3-v2-Private | consistent harness |
A | LMCA | consistent harness |
A | ARC-AGI-2 | consistent harness |
A | DTBench | consistent harness |
A | FrontierMath-Tier-4-v2-Private | consistent harness |
Top 10 models by rank
Frontier leaderboard
| # | Model | Lab | Leads | Avg gap | Evidence |
|---|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 8 | 7.6 | reproduced |
| 2 | GPT-6 Astra | OpenAI | 11 | 1.8 | unverified |
| 3 | GPT-5 | OpenAI | 4 | 28.4 | reproduced |
| 4 | GPT-5.5 | OpenAI | 4 | 13.5 | reproduced |
| 5 | Claude Fable 5.1 | Anthropic | 6 | 6.7 | unverified |
| 6 | Gemini 3 Pro | Google DeepMind | 3 | 17.1 | reproduced |
| 7 | Claude Opus 4.6 | Anthropic | 3 | 20.5 | reproduced |
| 8 | Claude Opus 5 | Anthropic | 2 | 8.2 | reproduced |
| 9 | Llama 3.1-405B | Meta AI | 2 | 37.1 | reproduced |
| 10 | GPT-4 (Mar 2023) | OpenAI | 2 | 32.6 | reproduced |
Leads = benchmarks topped; avg gap = mean points behind SOTA elsewhere. Evidence tags how each score was checked — a score is f(model, harness, protocol), not f(model).
Answer pages
What is the X benchmark?
Head-to-head
Popular matchups
Follow the record
Every new score, contradiction, and analysis as it lands — plain Atom, open in any feed reader. No account, no inbox.
Section map
Explore the record
- Leaderboard447
ranked by evidence
- Benchmarks62
graded on integrity
- Compare70
head-to-head matchups
- Labs37
developers tracked
- Explainers28
what each benchmark measures
- Papers60
source records
- Methodology
how we rank & grade