Skip to content

Early Adopter Offer:Get 40% off Core & Scale for your first year with code EARLYQFView pricing

AI in Testing

LLM-as-judge

Also known as: model-graded eval

LLM-as-judge is using one language model to grade another model output. It scales where human grading cannot, and it carries measured biases that have to be designed around rather than assumed away.

The biases are documented rather than hypothetical. The MT-Bench study found GPT-4 gave a consistent verdict when the two candidate answers were swapped only about 65% of the time, and that GPT-3.5 and Claude-v1 were fooled by a repetitive-list attack in 91.3% of cases. Supplying a reference answer cut grading failures on maths problems from 70% to 15%.

The standard mitigations follow from the biases: prefer pairwise comparison to open-ended scoring, randomise the position of candidate answers, provide reference answers where they exist, and calibrate the judge against human labels before trusting it. A frequently-cited self-preference effect is reported in the same study, which explicitly cautions that its data cannot establish whether the bias is real.

  • Position bias is large — verdict consistency under answer-swapping was ~65% for the strongest judge tested.
  • Reference-guided grading is the single highest-leverage fix.
  • Calibrate against human labels before gating anything on a judge.

See it in your own test results

Qualflare detects flaky tests, clusters failures by root cause, and scores release risk from the test results you already produce in CI. Start free.

Start free with Qualflare

← Back to the testing & observability glossary.

Last reviewed September 9, 2026