Skip to content

Edition·2026-09-20 · Sun

Fed from·Epoch AI·HF Open LLM·OpenAlex·OpenAI evals

The benchmark record for AI models.

Every reported score is tagged by how it was checked — independently reproduced, vendor self-reported, contradicted, or unverified — and every benchmark is graded on the four signals that decide whether a number still means something. We grade the evaluations, not just the models.

Latest measurement·2026-09-03·34 of 62 benchmarks graded A

At a glancesnapshot 2026-09-03
Benchmarks graded
62
34 rated A
Models tracked
447
Claims in ledger
2,859
Independently reproduced
23%
659 of 2,859

How trustworthy the evaluations are

Benchmark integrity

saturated · belief erodes50607080901000.20.40.60.81.0saturation — top score's proximity to ceiling →integrity →DTBenchOTIS Mock AIME 2024-2025GPQA diamondLAMBADA
ABCDdot size = models scored · 62 benchmarks

Trust holds the top-left; the bottom-right is where scores max out and belief erodes. The best-graded benchmarks:

A
Chess Puzzles
A
FrontierMath-Tiers-1-3-v2-Private
A
LMCA
A
ARC-AGI-2
A
DTBench
A
FrontierMath-Tier-4-v2-Private

Top 10 models by rank

Frontier leaderboard

Scores through 2026-09-03Full leaderboard →
#ModelLabLeadsAvg gapEvidence
1Claude Fable 5Anthropic87.6reproduced
2GPT-6 AstraOpenAI111.8unverified
3GPT-5OpenAI428.4reproduced
4GPT-5.5OpenAI413.5reproduced
5Claude Fable 5.1Anthropic66.7unverified
6Gemini 3 ProGoogle DeepMind317.1reproduced
7Claude Opus 4.6Anthropic320.5reproduced
8Claude Opus 5Anthropic28.2reproduced
9Llama 3.1-405BMeta AI237.1reproduced
10GPT-4 (Mar 2023)OpenAI232.6reproduced

Leads = benchmarks topped; avg gap = mean points behind SOTA elsewhere. Evidence tags how each score was checked — a score is f(model, harness, protocol), not f(model).

Answer pages

What is the X benchmark?

Head-to-head

Popular matchups

Follow the record

Every new score, contradiction, and analysis as it lands — plain Atom, open in any feed reader. No account, no inbox.

Section map

Explore the record