Skip to content

Early Adopter Offer:Get 40% off Core & Scale for your first year with code EARLYQFView pricing

AI in Testing

LLM eval

Also known as: eval, evaluation suite

An LLM eval is a scored test suite for a language-model application — a fixed set of inputs whose outputs are graded rather than compared for equality, because the same prompt does not reliably produce the same text.

Traditional assertions assume determinism. Language models do not provide it: even at temperature 0, providers document that identical inputs can produce different outputs, and one published experiment recorded 80 unique completions from 1,000 identical requests. An eval replaces the equality check with a grade, and the aggregate score becomes the thing you threshold.

Graders come in three classes with different trade-offs. Code-based graders — exact match on a classification, JSON schema validation, a tool-call assertion — are fast, cheap and reproducible but brittle to valid variation. Model-based graders capture nuance but are themselves non-deterministic and need calibrating against human judgement. Human grading is the gold standard and does not scale.

  • Scores output rather than asserting equality on it.
  • Push work to code-based graders wherever the task allows.
  • A single eval run is a sample, not a verdict — repeat and measure the distribution.

See it in your own test results

Qualflare detects flaky tests, clusters failures by root cause, and scores release risk from the test results you already produce in CI. Start free.

Start free with Qualflare

← Back to the testing & observability glossary.

Last reviewed September 9, 2026