Methodology · Reference scale 2026-Q4
How the llmrun Score is built
The llmrun Score rates a language model from 0 to 100, fitted to results from the public leaderboards we track. This page describes the calculation and its known limits.
- Reference benchmarks
- 40
- Sources in the score
- 4
- Frozen scale
- 2026-Q4
- Held-out error (RMSE)
- 0.093
01
Overview
llmrun does not run benchmarks. The score is computed from results that independent leaderboards have already published, and each source is credited.
- 1
Public leaderboards
Results published by 4 independent sources.
- 2
Normalize
Each result is rescaled to 0–1, from the benchmark's random-guess baseline to its ceiling.
- 3
IRT fit
One ability θ per model, plus a difficulty b and a discrimination a per benchmark.
- 4
Frozen scale
Reference 2026-Q4 fixes the parameters of 40 benchmarks, so scores don't drift.
- 5
llmrun Score
A 0–100 score with a 90% range, plus 5 skill scores.
02
Why not average the scores
Models are tested on different sets of benchmarks. A plain average favours a model run only on easy benchmarks over one submitted to hard ones, and averages taken over different test sets can't be compared with each other.
Model A
Weaker. Tested only on the easiest benchmarks.
- LiveBench Math81%
- MMLU-Pro86%
- LiveBench Reasoning73%
- MATH Level 594%
Model B
Stronger. Submitted only to the hardest benchmarks.
- Chess Puzzles31%
- FrontierCode35%
- Humanity's Last Exam33%
- CritPt16%
Plain average
Mean of whatever each model was tested on
- Model A
- 84
- Model B
- 29
Puts A far ahead, but the gap only reflects which tests each model took.
llmrun Score
Ability placed on one shared difficulty scale
- Model A
- 38.6
- Model B
- 62.4
Accounts for test difficulty and ranks B higher.
03
Item Response Theory
Standardized tests such as the SAT and GRE use Item Response Theory to compare test-takers who answered different questions. Epoch AI's Capabilities Index applies the same method to AI models. We use the two-parameter logistic (2PL) model.
Each model gets one ability θ. Each benchmark gets a difficulty b, the ability at which a model is expected to land halfway between random guessing and a perfect score, and a discrimination a, which sets how sharply the benchmark separates weaker models from stronger ones.
Before fitting, each result is normalized to 0–1 between the benchmark's random-guess baseline and its ceiling. On a four-option multiple-choice test, 25% counts as zero.
All models, open and proprietary, and all benchmarks are fitted together by maximum a posteriori estimation. Two models that never ran the same benchmark are still comparable, because each shares benchmarks with other models in the fit.
Expected normalized result
ŷ = σ(a · (θ − b))
- θ
- model ability
- b
- benchmark difficulty
- a
- benchmark discrimination
- σ
- logistic S-curve, 0 → 1
- MATH Level 5b -0.83 · a 2.65
- SWE-bench Verifiedb -0.13 · a 1.38
- Humanity's Last Examb 1.87 · a 1.03
- MATH Level 5
- 99%
- SWE-bench Verified
- 83%
- Humanity's Last Exam
- 29%
- llmrun Score
- 57.4
04
Which benchmarks count
The score is fitted on the active panel: benchmarks that new models are still being tested on.
Still in use
A benchmark counts while models released in the last 365 days are still submitted to it. Saturated older benchmarks such as GSM8K and HellaSwag drop out this way.
A fixed 0–100 scale
Relative ratings such as LMArena's Elo are excluded, because they have no baseline or ceiling to normalize against.
Plausible bounds
Benchmarks where more than 25% of results sit at exactly 0 or 100 are excluded. In that case the scale bounds are probably wrong.
Benchmark difficulty map
- Coding
- Agents
- Math
- Science
- Reasoning
- Overall only
- Two skills
- LiveBench Codingdifficulty -1.70
- LiveBench Mathdifficulty -1.57
- MMLU-Prodifficulty -1.05
- LiveBench Reasoningdifficulty -1.01
- MATH Level 5difficulty -0.83
- ForecastBenchdifficulty -0.82
- SWE-bench Multilingualdifficulty -0.46
- Fiction.LiveBenchdifficulty -0.37
- GPQA Diamonddifficulty -0.26
- DTBenchdifficulty -0.20
- SWE-bench Verifieddifficulty -0.13
- AIME 2024/2025difficulty -0.13
- Aider Polyglotdifficulty -0.09
- SWE-bench Bash Onlydifficulty 0.05
- WeirdMLdifficulty 0.37
- SWE-bench Litedifficulty 0.39
- ARC-AGIdifficulty 0.44
- SimpleBenchdifficulty 0.50
- Surface Evolver Benchdifficulty 0.59
- Terminal-Benchdifficulty 0.61
- FrontierMathdifficulty 0.81
- SciCodedifficulty 0.99
- ARC-AGI-2difficulty 1.05
- BALROGdifficulty 1.10
- DeepSWEdifficulty 1.11
- LMCAdifficulty 1.18
- ProofBenchdifficulty 1.20
- APEX-Agentsdifficulty 1.22
- SimpleQAdifficulty 1.23
- ALE-Benchdifficulty 1.25
- FrontierMath Tier 4difficulty 1.34
- Vending-Bench 2difficulty 1.59
- GSO-Benchdifficulty 1.73
- Mystery Game Puzzlesdifficulty 1.73
- Chess Puzzlesdifficulty 1.78
- FrontierCodedifficulty 1.79
- Humanity's Last Examdifficulty 1.87
- CritPtdifficulty 2.41
- CL-benchdifficulty 3.88
- CL-bench Lifedifficulty 4.38
All 40 reference parameters
| Benchmark | Skill | Difficulty b | Discrimination a | Source |
|---|---|---|---|---|
| LiveBench Coding | Coding | −1.70 | 0.46 | LiveBench |
| LiveBench Math | Math | −1.57 | 0.82 | LiveBench |
| MMLU-Pro | — | −1.05 | 1.44 | TIGER-Lab MMLU-Pro |
| LiveBench Reasoning | Reasoning | −1.01 | 0.83 | LiveBench |
| MATH Level 5 | Math | −0.83 | 2.65 | Epoch AI |
| ForecastBench | Reasoning | −0.82 | 0.31 | Epoch AI |
| SWE-bench Multilingual | Coding | −0.46 | 0.62 | SWE-bench |
| Fiction.LiveBench | — | −0.37 | 1.73 | Epoch AI |
| GPQA Diamond | Science | −0.26 | 1.67 | Epoch AI |
| DTBench | Reasoning | −0.20 | 1.37 | Epoch AI |
| SWE-bench Verified | Agents | −0.13 | 1.38 | SWE-bench |
| AIME 2024/2025 | Math | −0.13 | 2.67 | Epoch AI |
| Aider Polyglot | Coding | −0.09 | 2.13 | Epoch AI |
| SWE-bench Bash Only | Agents | +0.05 | 1.50 | SWE-bench |
| WeirdML | Coding | +0.37 | 1.15 | Epoch AI |
| SWE-bench Lite | Coding | +0.39 | 1.04 | SWE-bench |
| ARC-AGI | Reasoning | +0.44 | 2.91 | Epoch AI |
| SimpleBench | Reasoning | +0.50 | 0.90 | Epoch AI |
| Surface Evolver Bench | Coding | +0.59 | 1.69 | Epoch AI |
| Terminal-Bench | Agents | +0.61 | 1.65 | Epoch AI |
| FrontierMath | Math | +0.81 | 2.06 | Epoch AI |
| SciCode | Coding, Science | +0.99 | 0.51 | Epoch AI |
| ARC-AGI-2 | Reasoning | +1.05 | 3.64 | Epoch AI |
| BALROG | Agents | +1.10 | 0.68 | Epoch AI |
| DeepSWE | Agents | +1.11 | 1.57 | Epoch AI |
| LMCA | Reasoning | +1.18 | 0.90 | Epoch AI |
| ProofBench | Math | +1.20 | 3.45 | Epoch AI |
| APEX-Agents | Agents | +1.22 | 0.90 | Epoch AI |
| SimpleQA | — | +1.23 | 1.01 | Epoch AI |
| ALE-Bench | Coding | +1.25 | 1.30 | Epoch AI |
| FrontierMath Tier 4 | Math | +1.34 | 3.52 | Epoch AI |
| Vending-Bench 2 | Agents | +1.59 | 1.58 | Epoch AI |
| GSO-Bench | Agents | +1.73 | 1.59 | Epoch AI |
| Mystery Game Puzzles | Reasoning | +1.73 | 1.42 | Epoch AI |
| Chess Puzzles | Reasoning | +1.78 | 1.40 | Epoch AI |
| FrontierCode | Coding | +1.79 | 1.02 | Epoch AI |
| Humanity's Last Exam | Science | +1.87 | 1.03 | Epoch AI |
| CritPt | Science | +2.41 | 1.38 | Epoch AI |
| CL-bench | — | +3.88 | 0.46 | Epoch AI |
| CL-bench Life | — | +4.38 | 0.52 | Epoch AI |
05
From ability to a 0–100 score
Ability θ has no natural units, so we convert it. The llmrun Score is the model's expected normalized result, averaged over all 40 benchmarks of the frozen 2026-Q4 reference set and multiplied by 100.
Benchmarks a model was never run on still count in that average, using the fit's prediction for that model. This way every model is scored on the same set of benchmarks.
A score of 60 means the model is expected to get 60% of the way from random guessing to a perfect score on a typical current benchmark.
06
A fixed scale
If the fit were simply rerun on every data update, the scale would re-centre each time, and adding one very hard benchmark would move every model's number. So the scale is fixed.
Frozen reference
Difficulty and discrimination for the 40 reference benchmarks are fixed in reference 2026-Q4, created 2026-10-03.
New benchmarks are calibrated in
A benchmark added later is calibrated onto the existing scale. The scale itself does not move to fit it.
Scores follow a model's own results
Adding other models or new benchmarks does not move a model's score. It changes mainly when new results for that model arrive. Its range and skill scores can shift slightly as the fit is refined.
The current method is v3-irt-2 on reference 2026-Q4. If the scale is ever re-based, that will ship as a new version with a published old-to-new conversion table.
07
The 90% range
Each score is shown with a 90% range, computed as the score at θ ± 1.645 standard errors (Laplace approximation). Fewer or less informative benchmarks give a wider range.
- 62.490% range 58–66
Well covered
Many informative boards across all skills
- 62.490% range 44–77
Thinly covered
Meets the 4-benchmark minimum
08
When a score is published
A score based on a handful of results would look more precise than it is, so a model needs a minimum amount of evidence first.
- Overall llmrun Score
- At least 4 eligible benchmarks across at least 2 categories, one of which is reasoning or coding.
- Skill score
- At least 2 benchmarks in that skill. Otherwise the skill is marked as not enough data.
09
Skill scores
Each skill is a fixed set of benchmarks. A skill score is the model's expected result on that skill's reference benchmarks, on the same scale as the overall score. A benchmark can belong to two skills. SciCode, for example, counts for both coding and science.
Coding
9 benchmarksWriting and fixing code that must run and pass tests.
LiveBench Coding · Aider Polyglot · SciCode · WeirdML · SWE-bench Multilingual · SWE-bench Lite · FrontierCode · ALE-Bench · Surface Evolver Bench
Agents
8 benchmarksLong multi-step tasks in real environments: repositories, terminals, simulated businesses.
SWE-bench Verified · SWE-bench Bash Only · Terminal-Bench · DeepSWE · GSO-Bench · APEX-Agents · Vending-Bench 2 · BALROG
Math
6 benchmarksFrom competition problems to research-level proofs.
AIME 2024/2025 · MATH Level 5 · LiveBench Math · FrontierMath · FrontierMath Tier 4 · ProofBench
Science
4 benchmarksGraduate-level physics, chemistry and biology, plus scientific code.
GPQA Diamond · CritPt · SciCode · Humanity's Last Exam
Reasoning
9 benchmarksNovel puzzles, deduction, planning and forecasting.
ARC-AGI · ARC-AGI-2 · SimpleBench · LiveBench Reasoning · Chess Puzzles · DTBench · LMCA · Mystery Game Puzzles · ForecastBench
Planned
Knowledge, Long context, Instruction following, Multilingual, Vision. These will be added when enough open models have public results to rank them fairly.
10
Validation
We hold out part of the results, refit, and measure how well the fit predicts the held-out results (5-fold cross-validation). Error is RMSE on the 0–1 normalized scale, so lower is better.
Held-out prediction error (RMSE)
- llmrun Score (IRT)
- 0.093
- Model offset + board mean
- 0.177
- Board average only
- 0.246
11
llmrun Score vs Fit
The two ratings answer different questions and are computed separately.
llmrun Score · 0–100
How capable is the model?
Does not depend on hardware. The number is the same on a laptop and on a data-centre GPU.
Fit · S to F
How well does it run on my hardware?
VRAM headroom and speed on a specific GPU or device. How the S–F grades work
12
Limitations
Known gaps in the score and how we deal with them.
- Contamination
- A model trained on a benchmark's questions scores above its real ability. Benchmarks with fresh, rotating or private question sets (LiveBench, FrontierMath, ARC-AGI-2) reduce this effect but cannot remove it.
- Self-reported vs independent
- Some leaderboards accept runs submitted by developers, with their own scaffolds and settings. Independent evaluations are harder to game. We use each source's published figure as is.
- Effort and thinking variants
- Proprietary models are often listed at several reasoning-effort settings. We use the best variant of each, so its score reflects its strongest published configuration.
- Sparse coverage for small models
- Small open models are rarely submitted to frontier benchmarks, so many fall short of the minimum or get wide ranges. A missing score means there is not enough evidence. It does not mean the model is weak.
- English-centric
- Almost every reference benchmark is in English. Multilingual ability is not measured yet.
13
Sources
All results are published by these organisations. llmrun normalizes them and runs the IRT fit.
- Epoch AI Benchmarking Hub
Most of the panel: independent runs of frontier reasoning, math, science, coding and agent benchmarks (CC-BY).
32 in reference
- LiveBench
Coding, math and reasoning splits designed to resist contamination. New questions are added on a rolling basis.
3 in reference
- SWE-bench leaderboards
Official SWE-bench Verified, Bash Only, Lite and Multilingual results, scored on whether the model fixes real GitHub issues.
4 in reference
- TIGER-Lab MMLU-Pro
MMLU-Pro, a harder ten-choice successor to MMLU with more reasoning questions and wide open-model coverage.
1 in reference
- LMArena
Elo ratings from blind head-to-head human votes. Shown on separate leaderboards and not used in the score, because Elo is a relative rating, not a 0–100 result.
Not in score
Background on each benchmark is in the benchmark guide. The full ranking is on the llmrun Score leaderboard.
14
Frequently asked questions
- What is the llmrun Score?
- A 0–100 capability score for language models, computed from public benchmark results. It is the model's expected result averaged over the 40 benchmarks of the frozen 2026-Q4 reference set, where 0 is random guessing and 100 is a perfect score.
- Why not just average a model's benchmark scores?
- Each model is tested on its own mix of benchmarks, so a plain average mostly reflects how hard that mix was. The llmrun Score fits all models and all benchmarks together with Item Response Theory, which estimates each benchmark's difficulty and takes it into account.
- Does llmrun run the benchmarks itself?
- No. The score uses only results published by Epoch AI Benchmarking Hub, LiveBench, SWE-bench leaderboards, TIGER-Lab MMLU-Pro. Each source is credited and linked on this page.
- Why does every score come with a range?
- It is the 90% interval of the estimate. A model with results on many informative benchmarks gets a narrow range, and one that barely meets the minimum gets a wide one. If two ranges overlap heavily, treat the models as roughly tied.
- Will a model's score change when new models or harder benchmarks are added?
- No. Benchmark parameters are frozen in reference 2026-Q4, and new benchmarks are calibrated onto that scale. The overall score changes mainly when new results for that model come in, though its range and skill scores can shift slightly as the fit is refined.
- Why isn't LMArena part of the llmrun Score?
- LMArena publishes an Elo-style rating built from human preference votes. The rating only says how often voters preferred one model over another, so it has no fixed baseline or ceiling to normalize against. Arena ratings are shown on their own leaderboards.
- Why doesn't a model have an llmrun Score yet?
- An overall score needs results on at least 4 eligible benchmarks in at least 2 categories, including reasoning or coding, and a skill score needs at least 2 benchmarks in that skill. Many small and new open models have not been submitted to enough public leaderboards yet.
- How is the llmrun Score different from the S–F grade?
- The llmrun Score rates the model itself and does not depend on hardware. The S–F Fit grade rates how well the model runs on a given GPU or device, based on VRAM headroom and speed. A strong model can get an F on a small GPU, and a weaker one an S.