CURRICULUMREADING A BENCHMARK HONESTLY01

Why two labs report different scores for the same model

The same model, the same dataset, and a fifteen-point spread. Prompt format and scoring convention are not implementation details; they are the measurement.

Reading time
4 min
Words
920
Sources cited
4
Updated
2026-09-05
Prerequisites
None

Two decks, one model, two numbers

One slide says the model scores 63.6 on MMLU. Another says 48.8. Both are describing LLaMA-65B. Both ran the published MMLU dataset. Neither is lying, and neither number is wrong.

The Hugging Face Open LLM Leaderboard team took this apart in 2023 after users reported that leaderboard MMLU scores did not match the numbers in the LLaMA paper. Three widely used implementations were run against the same models and the same data.

MMLU accuracy for the LLaMA family under three implementations of the same benchmark.
Implementation65B30B13B7B
Original (Hendrycks et al.)0.6360.5840.4700.351
HELM0.6370.5830.4710.339
EleutherAI harness (Jan 2023)0.4880.4570.3770.342

The convention moved the score further than two model generations

Read down a column and the ranking is stable: bigger models score higher under every implementation. Read across the 65B row and the spread is 14.8 points. Under the original implementation, the entire gap between 65B and 30B is 5.2 points. The choice of measurement convention moved the number almost three times as far as doubling the model.

Two things differ between those implementations, and neither is a modelling decision. The first is the prompt. HELM and the EleutherAI harness prepend a Question: marker; the harness omits the subject line that the original implementation includes. The second is what gets scored. The original implementation and HELM compare the model's probabilities over the answer letters — A, B, C, D. The EleutherAI harness compares probabilities over the full answer strings.

That second difference is the large one. Scoring the letter asks which option the model prefers. Scoring the full string asks which continuation the model finds most likely, which also depends on how long each answer is and how it tokenises. They are different questions. Both get published under the label MMLU.

On some tasks the convention outweighs the model entirely

ARC-Challenge has two competing conventions in circulation. The original is a cloze task: the model sees Question: {question}\nAnswer: and the harness compares the likelihood of each completion. The alternative presents lettered options and scores the generated letter, as MMLU does.

Mistral-7B scores 50.1 (±2.86) under the cloze convention and 72.4 (±2.56) under the MMLU-style convention. That is a 22-point gap on one model, one dataset, one afternoon. Both are reported as ARC-Challenge accuracy.

The effect is systematic rather than anecdotal. Across models, switching MMLU from symbol scoring to cloze scoring costs 13 to 24 accuracy points, and perturbations no larger than reordering the answer choices move models by up to eight leaderboard positions.

One tokeniser flag, fifty-nine points

In February 2025 someone opened an issue against the EleutherAI harness reporting that add_bos_token was destabilising quantized Llama-3 evaluations. On OPEA/Llama-3.3-70B-Instruct-int3-sym-inc, arc_easy scored 0.2643 without the flag and 0.8523 with it.

0.2643 on a four-option task is the random-guessing baseline. The run without the flag measured nothing at all, and it would have been published as a quantization quality result. On the int2 build of the same model the flag moved MMLU from 0.7142 to 0.7606 and lambada_openai from 0.7013 to 0.7413.

add_bos_token decides whether a beginning-of-sequence token is prepended to the prompt. It is not a hyperparameter, an architecture choice or a decoding setting. It is a boolean about tokenisation, and on the wrong side of it a 70B model looks like a coin flip.

What a score has to carry to be a measurement

None of the above is a reason to stop running benchmarks. It is a reason to stop reporting a benchmark as a single number. A score is reproducible only if a reader can reconstruct the run, and reconstruction needs all of the following.

The minimum a reportable score carries

  • Harness name and exact commit or release version.
  • Task identifier and task version, because task definitions change under a stable name.
  • Number of few-shot examples, and whether they were fixed or sampled.
  • The prompt template, verbatim, including any chat template applied.
  • Scoring method: generated answer, letter log-probability, or full-string log-probability.
  • Tokeniser settings that affect the prompt, add_bos_token included.
  • Model revision — a commit SHA, not a name — and the precision or quantization it was served at.
  • Decoding parameters, or an explicit statement that scoring was log-probability based and did not sample.
  • The date of the run.

What this implies for a catalogue

This is the reason every record in the MOEModels catalogue is addressed by commit SHA rather than by model name. Two repositories that share a name and differ in revision, precision or quantization are different artifacts, and a score measured on one is not evidence about the other.

It is also the reason this project publishes no leaderboard. A ranking assembled from numbers produced under different conventions is not a ranking of models. It is a ranking of measurement choices.

The rule

A benchmark score without its harness configuration is not a measurement. It is an anecdote.

Every figure in this lesson is sourced or computed. Where a claim would not verify against its source, it was removed rather than qualified.