Methodology · Reference scale 2026-Q4

How the llmrun Score is built

The llmrun Score rates a language model from 0 to 100, fitted to results from the public leaderboards we track. This page describes the calculation and its known limits.

Reference benchmarks
40
Sources in the score
4
Frozen scale
2026-Q4
Held-out error (RMSE)
0.093

01

Overview

llmrun does not run benchmarks. The score is computed from results that independent leaderboards have already published, and each source is credited.

  1. 1

    Public leaderboards

    Results published by 4 independent sources.

  2. 2

    Normalize

    Each result is rescaled to 0–1, from the benchmark's random-guess baseline to its ceiling.

  3. 3

    IRT fit

    One ability θ per model, plus a difficulty b and a discrimination a per benchmark.

  4. 4

    Frozen scale

    Reference 2026-Q4 fixes the parameters of 40 benchmarks, so scores don't drift.

  5. 5

    llmrun Score

    A 0–100 score with a 90% range, plus 5 skill scores.

02

Why not average the scores

Models are tested on different sets of benchmarks. A plain average favours a model run only on easy benchmarks over one submitted to hard ones, and averages taken over different test sets can't be compared with each other.

Model A

Weaker. Tested only on the easiest benchmarks.

  • LiveBench Math81%
  • MMLU-Pro86%
  • LiveBench Reasoning73%
  • MATH Level 594%

Model B

Stronger. Submitted only to the hardest benchmarks.

  • Chess Puzzles31%
  • FrontierCode35%
  • Humanity's Last Exam33%
  • CritPt16%

Plain average

Mean of whatever each model was tested on

Model A
84
Model B
29

Puts A far ahead, but the gap only reflects which tests each model took.

llmrun Score

Ability placed on one shared difficulty scale

Model A
38.6
Model B
62.4

Accounts for test difficulty and ranks B higher.

Hypothetical models. Each result is what the fitted 2026-Q4 model predicts for a model of that ability on that benchmark, so the example uses each benchmark's real difficulty.

03

Item Response Theory

Standardized tests such as the SAT and GRE use Item Response Theory to compare test-takers who answered different questions. Epoch AI's Capabilities Index applies the same method to AI models. We use the two-parameter logistic (2PL) model.

Each model gets one ability θ. Each benchmark gets a difficulty b, the ability at which a model is expected to land halfway between random guessing and a perfect score, and a discrimination a, which sets how sharply the benchmark separates weaker models from stronger ones.

Before fitting, each result is normalized to 0–1 between the benchmark's random-guess baseline and its ceiling. On a four-option multiple-choice test, 25% counts as zero.

All models, open and proprietary, and all benchmarks are fitted together by maximum a posteriori estimation. Two models that never ran the same benchmark are still comparable, because each shares benchmarks with other models in the fit.

Expected normalized result

ŷ = σ(a · (θ − b))

θ
model ability
b
benchmark difficulty
a
benchmark discrimination
σ
logistic S-curve, 0 → 1
  • MATH Level 5b -0.83 · a 2.65
  • SWE-bench Verifiedb -0.13 · a 1.38
  • Humanity's Last Examb 1.87 · a 1.03
0%25%50%75%100%−3−2−10+1+2+3+4Model ability θ →
MATH Level 5
99%
SWE-bench Verified
83%
Humanity's Last Exam
29%
llmrun Score
57.4
Real parameters from the 2026-Q4 reference. Each curve crosses 50% at its difficulty b (open circles). A steeper curve, with a higher a, separates models more sharply near that point. Move the slider to see the expected result on each benchmark and the score for that ability.

04

Which benchmarks count

The score is fitted on the active panel: benchmarks that new models are still being tested on.

  • Still in use

    A benchmark counts while models released in the last 365 days are still submitted to it. Saturated older benchmarks such as GSM8K and HellaSwag drop out this way.

  • A fixed 0–100 scale

    Relative ratings such as LMArena's Elo are excluded, because they have no baseline or ceiling to normalize against.

  • Plausible bounds

    Benchmarks where more than 25% of results sit at exactly 0 or 100 are excluded. In that case the scale bounds are probably wrong.

Benchmark difficulty map

  • Coding
  • Agents
  • Math
  • Science
  • Reasoning
  • Overall only
  • Two skills
← easierharder →
  • LiveBench Codingdifficulty -1.70
  • LiveBench Mathdifficulty -1.57
  • MMLU-Prodifficulty -1.05
  • LiveBench Reasoningdifficulty -1.01
  • MATH Level 5difficulty -0.83
  • ForecastBenchdifficulty -0.82
  • SWE-bench Multilingualdifficulty -0.46
  • Fiction.LiveBenchdifficulty -0.37
  • GPQA Diamonddifficulty -0.26
  • DTBenchdifficulty -0.20
  • SWE-bench Verifieddifficulty -0.13
  • AIME 2024/2025difficulty -0.13
  • Aider Polyglotdifficulty -0.09
  • SWE-bench Bash Onlydifficulty 0.05
  • WeirdMLdifficulty 0.37
  • SWE-bench Litedifficulty 0.39
  • ARC-AGIdifficulty 0.44
  • SimpleBenchdifficulty 0.50
  • Surface Evolver Benchdifficulty 0.59
  • Terminal-Benchdifficulty 0.61
  • FrontierMathdifficulty 0.81
  • SciCodedifficulty 0.99
  • ARC-AGI-2difficulty 1.05
  • BALROGdifficulty 1.10
  • DeepSWEdifficulty 1.11
  • LMCAdifficulty 1.18
  • ProofBenchdifficulty 1.20
  • APEX-Agentsdifficulty 1.22
  • SimpleQAdifficulty 1.23
  • ALE-Benchdifficulty 1.25
  • FrontierMath Tier 4difficulty 1.34
  • Vending-Bench 2difficulty 1.59
  • GSO-Benchdifficulty 1.73
  • Mystery Game Puzzlesdifficulty 1.73
  • Chess Puzzlesdifficulty 1.78
  • FrontierCodedifficulty 1.79
  • Humanity's Last Examdifficulty 1.87
  • CritPtdifficulty 2.41
  • CL-benchdifficulty 3.88
  • CL-bench Lifedifficulty 4.38
All 40 reference parameters
BenchmarkSkillDifficulty bDiscrimination aSource
LiveBench CodingCoding−1.700.46LiveBench
LiveBench MathMath−1.570.82LiveBench
MMLU-Pro—−1.051.44TIGER-Lab MMLU-Pro
LiveBench ReasoningReasoning−1.010.83LiveBench
MATH Level 5Math−0.832.65Epoch AI
ForecastBenchReasoning−0.820.31Epoch AI
SWE-bench MultilingualCoding−0.460.62SWE-bench
Fiction.LiveBench—−0.371.73Epoch AI
GPQA DiamondScience−0.261.67Epoch AI
DTBenchReasoning−0.201.37Epoch AI
SWE-bench VerifiedAgents−0.131.38SWE-bench
AIME 2024/2025Math−0.132.67Epoch AI
Aider PolyglotCoding−0.092.13Epoch AI
SWE-bench Bash OnlyAgents+0.051.50SWE-bench
WeirdMLCoding+0.371.15Epoch AI
SWE-bench LiteCoding+0.391.04SWE-bench
ARC-AGIReasoning+0.442.91Epoch AI
SimpleBenchReasoning+0.500.90Epoch AI
Surface Evolver BenchCoding+0.591.69Epoch AI
Terminal-BenchAgents+0.611.65Epoch AI
FrontierMathMath+0.812.06Epoch AI
SciCodeCoding, Science+0.990.51Epoch AI
ARC-AGI-2Reasoning+1.053.64Epoch AI
BALROGAgents+1.100.68Epoch AI
DeepSWEAgents+1.111.57Epoch AI
LMCAReasoning+1.180.90Epoch AI
ProofBenchMath+1.203.45Epoch AI
APEX-AgentsAgents+1.220.90Epoch AI
SimpleQA—+1.231.01Epoch AI
ALE-BenchCoding+1.251.30Epoch AI
FrontierMath Tier 4Math+1.343.52Epoch AI
Vending-Bench 2Agents+1.591.58Epoch AI
GSO-BenchAgents+1.731.59Epoch AI
Mystery Game PuzzlesReasoning+1.731.42Epoch AI
Chess PuzzlesReasoning+1.781.40Epoch AI
FrontierCodeCoding+1.791.02Epoch AI
Humanity's Last ExamScience+1.871.03Epoch AI
CritPtScience+2.411.38Epoch AI
CL-bench—+3.880.46Epoch AI
CL-bench Life—+4.380.52Epoch AI
The 40 benchmarks of the 2026-Q4 reference, placed by fitted difficulty. 0 is the average model in the fit. Hover a benchmark to see its parameters and source.

05

From ability to a 0–100 score

Ability θ has no natural units, so we convert it. The llmrun Score is the model's expected normalized result, averaged over all 40 benchmarks of the frozen 2026-Q4 reference set and multiplied by 100.

Benchmarks a model was never run on still count in that average, using the fit's prediction for that model. This way every model is scored on the same set of benchmarks.

A score of 60 means the model is expected to get 60% of the way from random guessing to a perfect score on a typical current benchmark.

06

A fixed scale

If the fit were simply rerun on every data update, the scale would re-centre each time, and adding one very hard benchmark would move every model's number. So the scale is fixed.

Frozen reference

Difficulty and discrimination for the 40 reference benchmarks are fixed in reference 2026-Q4, created 2026-10-03.

New benchmarks are calibrated in

A benchmark added later is calibrated onto the existing scale. The scale itself does not move to fit it.

Scores follow a model's own results

Adding other models or new benchmarks does not move a model's score. It changes mainly when new results for that model arrive. Its range and skill scores can shift slightly as the fit is refined.

The current method is v3-irt-2 on reference 2026-Q4. If the scale is ever re-based, that will ship as a new version with a published old-to-new conversion table.

07

The 90% range

Each score is shown with a 90% range, computed as the score at θ ± 1.645 standard errors (Laplace approximation). Fewer or less informative benchmarks give a wider range.

Score
  • Well covered

    Many informative boards across all skills

    62.490% range 58–66
  • Thinly covered

    Meets the 4-benchmark minimum

    62.490% range 44–77
Two hypothetical models with the same estimated ability (θ = 1.2). The dot is the score and the band is its 90% range.

08

When a score is published

A score based on a handful of results would look more precise than it is, so a model needs a minimum amount of evidence first.

Overall llmrun Score
At least 4 eligible benchmarks across at least 2 categories, one of which is reasoning or coding.
Skill score
At least 2 benchmarks in that skill. Otherwise the skill is marked as not enough data.

09

Skill scores

Each skill is a fixed set of benchmarks. A skill score is the model's expected result on that skill's reference benchmarks, on the same scale as the overall score. A benchmark can belong to two skills. SciCode, for example, counts for both coding and science.

  • Coding

    9 benchmarks

    Writing and fixing code that must run and pass tests.

    LiveBench Coding · Aider Polyglot · SciCode · WeirdML · SWE-bench Multilingual · SWE-bench Lite · FrontierCode · ALE-Bench · Surface Evolver Bench

  • Agents

    8 benchmarks

    Long multi-step tasks in real environments: repositories, terminals, simulated businesses.

    SWE-bench Verified · SWE-bench Bash Only · Terminal-Bench · DeepSWE · GSO-Bench · APEX-Agents · Vending-Bench 2 · BALROG

  • Math

    6 benchmarks

    From competition problems to research-level proofs.

    AIME 2024/2025 · MATH Level 5 · LiveBench Math · FrontierMath · FrontierMath Tier 4 · ProofBench

  • Science

    4 benchmarks

    Graduate-level physics, chemistry and biology, plus scientific code.

    GPQA Diamond · CritPt · SciCode · Humanity's Last Exam

  • Reasoning

    9 benchmarks

    Novel puzzles, deduction, planning and forecasting.

    ARC-AGI · ARC-AGI-2 · SimpleBench · LiveBench Reasoning · Chess Puzzles · DTBench · LMCA · Mystery Game Puzzles · ForecastBench

  • Planned

    Knowledge, Long context, Instruction following, Multilingual, Vision. These will be added when enough open models have public results to rank them fairly.

10

Validation

We hold out part of the results, refit, and measure how well the fit predicts the held-out results (5-fold cross-validation). Error is RMSE on the 0–1 normalized scale, so lower is better.

Held-out prediction error (RMSE)

llmrun Score (IRT)
0.093
Model offset + board mean
0.177
Board average only
0.246
Removing any single large benchmark and refitting moves top-20 ranks by at most about 5 places.

11

llmrun Score vs Fit

The two ratings answer different questions and are computed separately.

llmrun Score · 0–100

How capable is the model?

Does not depend on hardware. The number is the same on a laptop and on a data-centre GPU.

Fit · S to F

How well does it run on my hardware?

VRAM headroom and speed on a specific GPU or device. How the S–F grades work

12

Limitations

Known gaps in the score and how we deal with them.

Contamination
A model trained on a benchmark's questions scores above its real ability. Benchmarks with fresh, rotating or private question sets (LiveBench, FrontierMath, ARC-AGI-2) reduce this effect but cannot remove it.
Self-reported vs independent
Some leaderboards accept runs submitted by developers, with their own scaffolds and settings. Independent evaluations are harder to game. We use each source's published figure as is.
Effort and thinking variants
Proprietary models are often listed at several reasoning-effort settings. We use the best variant of each, so its score reflects its strongest published configuration.
Sparse coverage for small models
Small open models are rarely submitted to frontier benchmarks, so many fall short of the minimum or get wide ranges. A missing score means there is not enough evidence. It does not mean the model is weak.
English-centric
Almost every reference benchmark is in English. Multilingual ability is not measured yet.

13

Sources

All results are published by these organisations. llmrun normalizes them and runs the IRT fit.

  • Epoch AI Benchmarking Hub

    Most of the panel: independent runs of frontier reasoning, math, science, coding and agent benchmarks (CC-BY).

    32 in reference

  • LiveBench

    Coding, math and reasoning splits designed to resist contamination. New questions are added on a rolling basis.

    3 in reference

  • SWE-bench leaderboards

    Official SWE-bench Verified, Bash Only, Lite and Multilingual results, scored on whether the model fixes real GitHub issues.

    4 in reference

  • TIGER-Lab MMLU-Pro

    MMLU-Pro, a harder ten-choice successor to MMLU with more reasoning questions and wide open-model coverage.

    1 in reference

  • LMArena

    Elo ratings from blind head-to-head human votes. Shown on separate leaderboards and not used in the score, because Elo is a relative rating, not a 0–100 result.

    Not in score

Background on each benchmark is in the benchmark guide. The full ranking is on the llmrun Score leaderboard.

14

Frequently asked questions

What is the llmrun Score?
A 0–100 capability score for language models, computed from public benchmark results. It is the model's expected result averaged over the 40 benchmarks of the frozen 2026-Q4 reference set, where 0 is random guessing and 100 is a perfect score.
Why not just average a model's benchmark scores?
Each model is tested on its own mix of benchmarks, so a plain average mostly reflects how hard that mix was. The llmrun Score fits all models and all benchmarks together with Item Response Theory, which estimates each benchmark's difficulty and takes it into account.
Does llmrun run the benchmarks itself?
No. The score uses only results published by Epoch AI Benchmarking Hub, LiveBench, SWE-bench leaderboards, TIGER-Lab MMLU-Pro. Each source is credited and linked on this page.
Why does every score come with a range?
It is the 90% interval of the estimate. A model with results on many informative benchmarks gets a narrow range, and one that barely meets the minimum gets a wide one. If two ranges overlap heavily, treat the models as roughly tied.
Will a model's score change when new models or harder benchmarks are added?
No. Benchmark parameters are frozen in reference 2026-Q4, and new benchmarks are calibrated onto that scale. The overall score changes mainly when new results for that model come in, though its range and skill scores can shift slightly as the fit is refined.
Why isn't LMArena part of the llmrun Score?
LMArena publishes an Elo-style rating built from human preference votes. The rating only says how often voters preferred one model over another, so it has no fixed baseline or ceiling to normalize against. Arena ratings are shown on their own leaderboards.
Why doesn't a model have an llmrun Score yet?
An overall score needs results on at least 4 eligible benchmarks in at least 2 categories, including reasoning or coding, and a skill score needs at least 2 benchmarks in that skill. Many small and new open models have not been submitted to enough public leaderboards yet.
How is the llmrun Score different from the S–F grade?
The llmrun Score rates the model itself and does not depend on hardware. The S–F Fit grade rates how well the model runs on a given GPU or device, based on VRAM headroom and speed. A strong model can get an F on a small GPU, and a weaker one an S.