A two-point win on a 198-question exam
A model card reports a GPQA Diamond score two points above a competitor's, and the launch post calls it a lead. GPQA Diamond contains 198 questions. Two points is four questions, and four questions is well inside the noise of which questions happened to get written.
This is not a hard problem. It is a standard two-proportion comparison, the formula is a century old, and the reason it is rarely applied to language-model evaluations is that the result is usually inconvenient.
The size the comparison actually needs
Evan Miller's Anthropic paper on eval statistics works the planning problem directly. Its worked figure is that detecting a three-percentage-point difference at 80% power requires 969 questions, and its general recommendation is that a new eval should carry at least 1,000 questions to have good signalling ability.
969 gets quoted as though it were a constant. It is not. It is the paired case, and the requirement depends entirely on the correlation between the two models' per-question outcomes. Run the same formula at a 50% baseline with the two runs treated as independent and the requirement is 4,361 questions; at a correlation of about 0.78 it is 969; at a correlation of 0.5 it is 2,181. Those three figures come from the site's own sample-size function, which is the standard two-proportion formula and nothing more.
The practical reading is that the paired design is worth roughly a four-fold reduction in the questions you need, and that you only get it if you can score both models on the same items.
GPQA Diamond has 198 questions. A three-point claim on it is roughly five times short of the paired requirement and twenty-two times short of the unpaired one. AIME has thirty questions per paper, where a single question moves the score 3.3 points. HumanEval has 164.
None of this means small benchmarks are useless. It means a small benchmark can support a large claim and cannot support a small one, and the boundary between the two is computable in advance.
The floor, in one formula
Scoring a benchmark is a sequence of Bernoulli trials, so the standard error of the accuracy is the square root of p(1−p)/n. Variance peaks at p = 0.5, which is the conservative case and the one worth quoting. The table below computes it live for published benchmark sizes.
Two cautions. First, this is a floor: it assumes questions are independent and that each question was answered once. Second, the interval it produces is for a single score. Comparing two scores widens it, which the next two sections address.
What each benchmark size can resolve
Computed at 50% accuracy, where the variance of a proportion is at its maximum. The last column is the smallest difference between two independent runs that a 95% interval would exclude zero for. Anything below it is inside the noise.
| Benchmark | Questions | Standard error | 95% margin, one score | Smallest resolvable gap |
|---|---|---|---|---|
| GPQA Diamond | 198 | 3.55 pp | ±6.96 pp | 9.85 pp |
| AIME 2025 | 30 | 9.13 pp | ±17.89 pp | 25.30 pp |
| MATH-500 | 500 | 2.24 pp | ±4.38 pp | 6.20 pp |
| HumanEval | 164 | 3.90 pp | ±7.65 pp | 10.82 pp |
| MMLU | 14,042 | 0.42 pp | ±0.83 pp | 1.17 pp |
| MMLU-Pro | 12,032 | 0.46 pp | ±0.89 pp | 1.26 pp |
| SWE-bench Verified | 500 | 2.24 pp | ±4.38 pp | 6.20 pp |
Sizes are sourced per benchmark and linked above. The arithmetic is the standard two-proportion framework, and it runs on the same module as the rest of the site.
Compare on the same questions, not on two averages
The common mistake is to subtract two published averages and treat the difference as a measurement. That throws away the strongest structure in the data. Models fail on the same hard questions, so their per-question outcomes are positively correlated.
The variance of a difference is Var(A) + Var(B) − 2·ρ·SE(A)·SE(B). When ρ is positive, the paired comparison is strictly tighter than the unpaired one, at no cost. Miller's framing is blunt: because eval question scores are likely to be positively correlated even across unrelated models, paired differences are a free reduction in estimator variance.
Getting that reduction requires per-question outputs from both models on the same questions. If a vendor publishes only an average, the paired comparison is not available to you, and the honest interval is the wide one.
Grouped questions break the independence assumption
Many benchmarks are not a bag of independent items. Reading-comprehension evals ask several questions about one passage. SWE-bench Verified instances cluster by repository. Once items are grouped, one lucky passage moves several scores together, and the naive standard error understates the real uncertainty.
The correction is a design effect: multiply the variance by 1 + (m − 1)·ICC, where m is the average cluster size and ICC is the intra-cluster correlation. Measured on two popular evals with Anthropic models, clustered standard errors ran from 1.10× to 3.05× the naive figure.
A three-fold error bar is not a rounding difference. It is the difference between a result and a coin flip.
Resampling reduces the variance you control
When a model is sampled rather than scored by log-probability, each question carries its own within-question variance. Answering every question K times drives that component down as σ²/K.
The paper's illustration: moving from one sample per question to two removes about a third of the variance, and six samples about five ninths. It never reaches zero, because the between-question variance — the fact that someone chose these 198 questions and not 198 others — is untouched by resampling. More samples cannot fix too few questions.
What to report
A comparison is publishable when a reader can decide, without trusting you, whether the difference is real.
The reportable form of a model comparison
- Both scores, the number of questions, and the standard error of each.
- The paired per-question difference with its own standard error, or an explicit statement that per-question outputs were unavailable.
- Clustered standard errors wherever questions are grouped, with the cluster definition named.
- The number of samples per question, and the decoding parameters used to draw them.
- The smallest difference the benchmark could have resolved at this size — computed before the run, not after.
- A statement of what the interval excludes, rather than a claim about which model is better.
The rule
State the smallest difference the benchmark could resolve before you report the difference you found. If the second is smaller than the first, you have measured the benchmark, not the models.