Skip to content

Early Adopter Offer:Get 40% off Core & Scale for your first year with code EARLYQFView pricing

Test Observability

Test baseline

Also known as: baseline run, reference run

A test baseline is a stored reference point — a previous run, a branch, or a recorded set of expected results — that a new run is compared against, turning an absolute result into a delta.

Most useful testing questions are comparative rather than absolute. A p95 latency of 190ms means little; 190ms against last week’s 135ms is a regression. A failing test means little; a test that passed on the target branch and fails on yours is a specific, actionable signal.

Baselines appear across testing under different names: a target-branch comparison in a merge request, a starred run in a load-testing tool, a stored screenshot in visual testing, a golden dataset in an eval suite. All of them are the same move — fixing a reference so a result can be read as a change.

  • Turns an absolute result into a delta, which is usually the useful form.
  • Requires storing runs; a single execution has nothing to compare against.
  • Same idea appears as target-branch diffs, starred runs, screenshots and golden datasets.

See it in your own test results

Qualflare detects flaky tests, clusters failures by root cause, and scores release risk from the test results you already produce in CI. Start free.

Start free with Qualflare

← Back to the testing & observability glossary.

Last reviewed September 9, 2026