CURRICULUMREADING A BENCHMARK HONESTLY04

LLM-as-judge and its failure modes

A judge that agrees with itself a quarter of the time when you swap the answers is not measuring quality. It is measuring position.

Reading time
4 min
Words
940
Sources cited
2
Updated
2026-09-05
Prerequisites
How many questions does it take to tell two models apart

The judge preferred the answer on the left

The MT-Bench authors ran the obvious control: present a judge with two answers, record the verdict, swap the order, ask again. A judge measuring quality gives the same verdict both times.

With the default prompt, GPT-4 was consistent 65.0% of the time, GPT-3.5 46.2%, and Claude-v1 23.8%. Only GPT-4 cleared 60%.

A pipeline that scores each pair once, in a fixed order, and uses Claude-v1 as the judge is letting position decide a large share of its results. That pipeline still produces a leaderboard. The leaderboard is partly a function of which column each system was written into.

Verbosity is a reward-hacking surface

The same paper tested a deliberately crude attack: rewrite an answer as a repetitive list that adds no information, only length. The attack succeeded against GPT-3.5 91.3% of the time and against Claude-v1 91.3% of the time. GPT-4 fell for it 8.7% of the time.

This matters more than it first appears, because a judge is rarely used only for reporting. As soon as a judge scores a training or selection loop, anything the judge rewards becomes a gradient. A 91% success rate for a hand-written attack is a lower bound on what an optimiser will find.

Self-preference tracks self-recognition

MT-Bench observed GPT-4 favouring its own answers by about ten points of win rate and Claude-v1 by about twenty-five, and then declined to conclude anything from it — limited data, small differences. That restraint was correct at the time.

The mechanism was established the following year. Panickssery and colleagues showed that LLM evaluators recognise their own generations, and that fine-tuning a model to change its self-recognition accuracy moves its self-preference in a linear relationship. The bias is not a quirk of one release; it is coupled to a capability that improves as models improve.

The operational consequence: a model must never be a judge in an evaluation where it, or a close sibling, is a candidate. That is a structural conflict, not a hygiene preference.

What a validated judge actually achieves

The case for judges is empirical and it is strong. On MT-Bench, GPT-4's pairwise agreement with human experts reaches 85%, which is higher than the 81% agreement humans reached with each other. On the crowdsourced Chatbot Arena data, GPT-4 single-answer grading reached 85% against 87% human-to-human agreement.

Read the qualifier attached to those numbers. They are measured under setup S2, without ties — that is, with the cases where a judge or a human declined to pick a winner excluded from the denominator.

Ties are exactly the close comparisons. Excluding them removes the region where the judge is least reliable, which is also the region where you usually need an answer, because two systems that differ obviously do not require a judge to separate them.

So 85% is a ceiling reported for the strongest available judge on a benchmark designed for judging, with the hardest cases set aside. It is not a default you inherit by calling an API. Your judge's agreement with your humans on your task is an unknown until you measure it.

The mitigation protocol

Judges remain useful. They are cheap, they scale, and for coarse discrimination they agree with humans well enough to be worth running. They stop being usable as the sole evidence for a small margin between two systems. The following is the minimum protocol that makes a judged result defensible.

  1. Randomise presentation order and score every pair in both orders. Count a disagreement between the two orders as a tie, and report the tie rate — it is your position-bias measurement.
  2. Run two judges from different model families. Report their agreement rate. Divergence is information, not an inconvenience to be averaged away.
  3. Control for length. Either report the length distribution of each candidate's answers alongside the win rate, or match on length before comparing.
  4. Hold out a human-labelled calibration set and publish judge-human agreement on it. A judge with no reported human agreement is an unvalidated instrument.
  5. Never judge with a model that is, or is derived from, a candidate.
  6. Publish the judge's model version, prompt verbatim, and decoding parameters, exactly as you would for a benchmark harness.

What a judged number is worth

With the protocol above, a judged comparison can be written in a form a reader can check. The shape of the sentence, with your own measured values substituted, is: across N prompts, two independent judges preferred A over B in x% and y% of pairs, at an order-disagreement rate of d%, with judge-human agreement of a% on a held-out calibration set of k items.

Without the protocol, a judged comparison supports only the claim that a model preferred one of these. The statistics from the previous lesson apply unchanged either way — a judged win rate is still a proportion over a finite sample, so it still needs an error bar, and the sample is still the prompts you chose.

The rule

Report the judge's order-disagreement rate and its agreement with humans, or do not report the judge's verdict as a result.

Every figure in this lesson is sourced or computed. Where a claim would not verify against its source, it was removed rather than qualified.