Testing & observability glossary
Plain-English definitions of the terms behind modern test automation, CI, and test observability — flaky tests, quality gates, test impact analysis, failure clustering, DORA metrics, and more. Each entry answers the question first, then goes deeper.
Flaky Tests & Reliability
-
Flaky test
A flaky test is a test that produces different results — sometimes passing, sometimes failing — on the same code, without any change to that code.
-
Flaky test detection
Flaky test detection is the practice of identifying tests that fail intermittently by analyzing their pass/fail history across many runs, rather than from a single result.
-
Non-determinism (in tests)
Non-determinism in tests is when the same test and the same code can yield different outcomes because the result depends on uncontrolled factors like timing, ordering, or shared state.
-
Test quarantine
Test quarantine is the practice of moving known-flaky tests out of the blocking path of a build so they stop failing the pipeline while their result is still recorded for triage.
-
Test retry
A test retry automatically re-runs a failed test a set number of times and passes it if any attempt succeeds — a way to absorb flakiness so it doesn’t fail the build.
Test Observability
-
CTRF
CTRF (Common Test Report Format) is a JSON schema for test results that treats retries, flakiness and per-attempt history as first-class fields rather than improvised conventions.
-
Flake rate
Flake rate is the percentage of test runs (or failures) that are flaky rather than genuine — a headline metric for how much you can trust your test suite.
-
JUnit XML
JUnit XML is the de facto standard file format for test results in CI — an XML document of testsuite and testcase elements that, despite universal support, has no official specification.
-
Mean time to detection (MTTD)
Mean time to detection (MTTD) is the average time between a defect being introduced and a test or signal catching it — a measure of how fast your quality feedback loop is.
-
Test artifact
A test artifact is a file a CI run produces and stores — a results file, an HTML report, a screenshot, a trace — retained under a policy rather than kept indefinitely.
-
Test baseline
A test baseline is a stored reference point — a previous run, a branch, or a recorded set of expected results — that a new run is compared against, turning an absolute result into a delta.
-
Test observability
Test observability is the ability to understand why your tests pass or fail over time by collecting and analyzing test results across every run — not just whether a single run was green.
AI in Testing
-
Agentic testing
Agentic testing is the use of AI agents that plan and carry out testing tasks with limited human direction — generating tests, running them, diagnosing failures, and adapting — rather than only assisting a human one step at a time.
-
Failure clustering
Failure clustering is the automatic grouping of many test failures that share the same underlying cause, so a wall of red collapses into a handful of distinct problems to fix.
-
LLM eval
An LLM eval is a scored test suite for a language-model application — a fixed set of inputs whose outputs are graded rather than compared for equality, because the same prompt does not reliably produce the same text.
-
LLM-as-judge
LLM-as-judge is using one language model to grade another model output. It scales where human grading cannot, and it carries measured biases that have to be designed around rather than assumed away.
-
Predictive flaky scoring
Predictive flaky scoring uses a test’s historical behavior to assign it a flakiness probability, flagging unreliable tests before they block a release rather than after.
-
Self-healing tests
Self-healing tests automatically repair their own locators or steps when the application changes — for example when a selector moves or is renamed — so a cosmetic UI change does not break the test.
-
Smart test selection
Smart test selection runs only the tests most likely to be affected by a given code change, instead of the entire suite, to cut CI time while preserving the chance of catching regressions.
CI/CD & Velocity
-
CI feedback loop
The CI feedback loop is the time between pushing a code change and getting a usable test result back. The shorter it is, the sooner developers can act while the change is still fresh in their heads.
-
DORA metrics
DORA metrics are four research-backed measures of software delivery performance: deployment frequency, lead time for changes, change failure rate, and time to restore service.
-
Monorepo testing
Monorepo testing is running and aggregating tests for many projects or packages that live in one repository, where a single commit can touch code shared across several of them.
-
Quality gate
A quality gate is an automated pass/fail checkpoint in a CI/CD pipeline that blocks a build from progressing unless it meets defined criteria — like test pass rate, coverage, or flakiness thresholds.
-
Test impact analysis (TIA)
Test impact analysis (TIA) determines which tests are affected by a specific code change so CI can run just those tests instead of the whole suite.
-
Test parallelization
Test parallelization runs multiple tests at the same time — across threads, processes, or machines — instead of one after another, to shorten the total run.
-
Test sharding
Test sharding splits a test suite into multiple subsets (shards) that run on separate machines or CI jobs at the same time, cutting total wall-clock time.
Mobile Testing
-
.xcresult bundle
An .xcresult bundle is the result container Xcode and xcodebuild produce for a test run — a directory holding test outcomes, logs, attachments, screenshots and coverage in Apple’s own format rather than in JUnit XML.
-
Accessibility identifier
An accessibility identifier is a stable, developer-assigned string attached to a UI element so tests can locate it without depending on its visible label, position, or view hierarchy.
-
Crash-free session rate
Crash-free session rate is the percentage of app sessions that finish without a crash — the headline stability metric mobile teams gate releases on, reported by crash-reporting SDKs rather than by the test suite.
-
Device fragmentation
Device fragmentation is the spread of hardware models, screen sizes, OS versions and vendor customisations a mobile app must run on — the reason a mobile test matrix is combinatorial rather than a single environment.
-
IdlingResource
An IdlingResource is an Espresso interface that tells the test runner when a background operation has finished, extending Espresso’s automatic synchronisation to work the main thread cannot see.
-
Instrumentation test
An instrumentation test is an Android test that runs on a device or emulator with access to the real Android framework, as opposed to a local unit test that runs on the JVM of the build machine.
-
Real device testing
Real device testing runs mobile tests on physical hardware rather than on an emulator or simulator, catching the class of defects that only appear with real sensors, real GPUs, real networks and real thermal behaviour.
-
Staged rollout
A staged rollout releases a mobile app update to a growing percentage of users over days rather than to everyone at once, so field stability can be observed before full exposure.
QA Foundations
-
Defect escape rate
Defect escape rate is the share of defects that reach production rather than being caught by testing — the most direct measure of whether a test suite is doing its job.
-
End-to-end testing
End-to-end testing exercises a complete user journey through a running system — browser or app, real services, real database — verifying that the integrated whole behaves correctly rather than any single component.
-
Gherkin
Gherkin is a plain-language syntax for writing executable test scenarios in a Given/When/Then structure, bridging business requirements and automated tests in BDD.
-
Monkey testing
Monkey testing feeds random inputs and actions to software to see if it crashes or misbehaves, surfacing edge cases and stability issues that scripted tests miss.
-
Regression testing
Regression testing re-runs existing tests to confirm that a change has not broken behaviour that previously worked — the reason most automated suites exist at all.
-
Release readiness
Release readiness is an assessment of whether a build is safe to ship, based on signals like test pass rate, flakiness, failure clusters, coverage of critical paths, and open risks.
-
Shift-left testing
Shift-left testing means moving testing earlier in the development process — closer to when code is written — so defects are caught sooner and cost less to fix.
-
Smoke testing
Smoke testing is a quick, shallow check that the critical functions of a build work, verifying it is stable enough for deeper testing to begin.
-
System integration testing
System integration testing (SIT) verifies that separately built modules or systems work correctly together, focusing on the interfaces and data flow between components.
-
Test debt
Test debt is the accumulated cost of neglected test suites — flaky tests, brittle selectors, poor coverage, and slow runs — that makes testing progressively harder and less trustworthy over time.
-
Test pyramid
The test pyramid is a strategy that favors many fast, cheap unit tests at the base, fewer integration tests in the middle, and a small number of slow end-to-end tests at the top.
-
Unit testing
Unit testing verifies a single piece of code — usually one function or class — in isolation from its dependencies, which makes unit tests the fastest and most stable layer of a suite.
-
User acceptance testing
User acceptance testing (UAT) is the final phase where real users verify that software meets business requirements and works in real-world scenarios before release.
-
Visual regression testing
Visual regression testing compares a rendered screenshot against a stored baseline image to catch unintended UI changes that assertion-based tests would miss.
Put these concepts to work
Qualflare turns your CI test results into AI failure clustering, flaky-test detection, and release-risk scoring. Start free — get your first analysis in minutes.
Start free with Qualflare