Educational guide
Quality gates in testing
What a quality gate actually decides, why code gates and test gates answer different questions, the rules teams use, and how to add one without training everyone to ignore it.
TL;DR
- A quality gate is an automated checkpoint. The next stage runs only when your conditions hold.
- Most tools gate on static code analysis. Gating on test execution results is the part teams skip.
- Start with one rule. A gate nobody trusts gets bypassed within a month.
- Budget a failure allowance, or flaky tests will teach everyone to ignore the gate.
A gate is only as honest as the run behind it. A suite with unexecuted cases can report green for work nobody did. Check completion before you check pass rate.
What is a quality gate?
A quality gate is an automated checkpoint in a delivery pipeline that lets the next stage run only when defined conditions are met. It takes a judgement that used to happen in a meeting, is this good enough to ship, and turns it into a rule a machine applies on every commit. The gate reads a signal, compares it against a threshold you set, and returns a verdict the pipeline acts on. In practice that verdict is a process exit code: zero lets the build continue, non-zero stops it. The point is not to catch everything. The point is to make one specific condition non-negotiable, so it stops depending on whether anyone remembered to look.
- Signal
- The evidence the gate reads. Coverage numbers, static analysis findings, or the results of an executed test run.
- Rule
- The threshold you set. No failed cases, pass rate above 95 percent, nothing left unexecuted.
- Verdict
- Pass or fail, expressed as an exit code the runner already knows how to interpret.
- Consequence
- What the pipeline does with the verdict. Stop the build, block the merge, hold the deploy.
A report tells you what happened and leaves the decision to a human. A gate makes the decision and removes the option to proceed. If someone can carry on past it without a conversation, what you have is a report.
Code quality gates and test quality gates are not the same thing
Search for quality gates and nearly everything you find describes static analysis: coverage percentage, duplicated lines, security hotspots, new issues on changed code. Those gates are useful, and they answer a narrow question, which is whether the code is well formed. They cannot tell you whether the software did what it was supposed to do, because they never ran it. A test quality gate reads execution results instead: which cases ran, which passed, which are still waiting on a human. Many teams have the first kind and assume they are covered.
| Code quality gate | Test quality gate | |
|---|---|---|
| Reads | Source code, without running it | Results of an executed test run |
| Answers | Is this code well formed? | Did the software behave correctly? |
| Typical rule | Coverage above 80 percent, no new critical issues | No failed cases, pass rate above 95 percent |
| Catches | Untested paths, risky patterns, known bad constructs | Regressions, broken flows, environment faults |
| Misses | Anything only visible at runtime | Anything no test covers |
| Runs | On every commit, in seconds | After a test run, which may include manual work |
They fail differently, so they catch different things. Coverage above 80 percent says nothing about whether the covered paths passed. A green test run says nothing about the paths nobody wrote a test for.
The rules teams actually use
Every gate is one rule, or a small stack of them, and each rule you add narrows what can get through. Start with one. Turning on all five at once produces a gate that fails for reasons nobody can explain, which is the fastest route to having it disabled.
| Rule | What it blocks | Use it when |
|---|---|---|
| Any failure blocks | The run contains at least one failed case | The suite is stable and a failure is always real |
| Widen the statuses | Failed or blocked cases both stop the build | A blocked case means the same thing to you as a failure |
| Failure allowance | More than N failures, but not N or fewer | You have a known-flaky tail you are actively shrinking |
| Minimum pass rate | Pass rate below a percentage of what executed | Suite size varies a lot between runs |
| Require completion | Any case still unexecuted | The run includes manual sign-off, or partial runs are common |
A failure allowance of three is a promise to fix three tests. Write down the number and the date you expect it to reach zero. Otherwise it becomes permanent and the gate quietly stops meaning anything.
Choosing your first rule
- Pull the last 20 runs of the suite you want to gate. Count how many were red, and how many of those were real defects.
- If most red runs were real defects, start with any failure blocks.
- If most were flake, start with require completion on its own. It is honest, it never fires falsely, and it buys you time to fix the suite.
- Add a second rule only after the first has run for a month without anyone asking to bypass it.
Where the gate goes
The same rule behaves very differently depending on where in the pipeline it sits. A gate before merge protects the main branch and has to be fast. A gate before deploy protects users and can afford to wait. Putting a slow gate in a fast position is the most common way teams end up removing it.
Before merge
Runs on the pull request. Has to finish in minutes, so it gates on the automated suite only. Keeps a regression from entering the main branch, where it becomes everyone's problem.
Before deploy to staging
Runs after the build. Can include a longer automated suite. This is where most teams get the best value for the least friction.
Before deploy to production
Runs after the release candidate is built. Can wait on manual sign-off and exploratory testing, because the cost of shipping a defect here is highest.
Automated results arrive in seconds. Manual sign-off does not. A production gate that can poll a run until every case has been executed lets a deploy stage wait for QA instead of racing ahead of it. Without that, the gate reads a half-finished run and passes it.
Gating on work people do
Every gate described so far reads a run that a machine finished. The moment a person is part of the run, the gate gains a new problem: the results are not there yet, and "not there yet" looks exactly like "nothing failed". A gate that cannot tell those two apart will happily pass a release on the day QA started testing it. Two behaviours separate them, and which one you want depends on whether the pipeline should fail or wait.
Fail now: the results should already exist
The pipeline expects a finished run. If anything is still unexecuted when the gate fires, that is a fault worth stopping for rather than something to wait out. Use this before merge, and on any stage where a person was never in the loop.
Wait, then decide: someone is still working
The pipeline expects to arrive before the humans finish. The gate polls the run on an interval and evaluates only once nothing is left unexecuted, so the deploy stage holds instead of racing ahead. Use this before a release that needs sign-off.
Take a run of 60 cases where 20 have passed and 40 are untouched. It has zero failures, a 100 percent pass rate, and it clears a failure allowance of any size. Every rule except completion reads it as green. This is the most common way a manual gate gets skipped, and it fails silently, which is worse than failing loudly.
| Run state when the gate fires | Fail now | Wait, then decide |
|---|---|---|
| Everything executed, all passed | Passes | Passes |
| Everything executed, some failed | Fails | Fails |
| Some cases still unexecuted | Fails straight away | Holds, re-checks, then evaluates |
| Still unexecuted when the time runs out | Not applicable | Fails |
| Nobody has started the run | Fails straight away | Holds for the full timeout, then fails |
Setting the wait
- Time the last five manual runs of this plan from start to sign-off. Take the longest, not the average.
- Set the timeout above that, with headroom for a slow day. Thirty minutes suits a smoke pass; a full regression cycle may need hours.
- Pick a poll interval short enough to add no meaningful delay and long enough to be polite to the API. Fifteen to thirty seconds is normal.
- Decide what expiry means before you need it. A gate that times out has learned nothing, so it has to fail. Treating a timeout as a pass reopens exactly the hole the gate was closing.
A gate that waits is not policing testers. It removes a coordination step that used to happen in chat, where somebody asks whether QA has finished and somebody else answers from memory. The pipeline asks continuously and acts on the answer, so nobody has to remember to.
What it looks like in a pipeline
A gate is one step. It runs after your results exist, reads them, and returns an exit code. Every major CI system already fails the surrounding step on a non-zero exit, so no special plugin is needed to make this work.
# 1. run the tests
run-tests --output results.xml
# 2. publish the results where the gate can read them
publish-results --file results.xml
# 3. read them back and decide
check-gate --fail-on failed --require-complete
# exit 0 -> pipeline continues
# exit 1 -> pipeline stops here
# exit 2 -> the gate itself broke, treat as a failure- Exit codes are the contract. Zero means proceed, non-zero means stop. Reserve a distinct code for usage and API errors so a broken gate is never mistaken for a failing one.
- Read results live, not from a cache. A gate that reads a summary computed earlier will happily pass a run whose results changed seconds ago.
- Log the numbers it saw. A gate that prints only pass or fail sends whoever it blocked off hunting. Print the counts and the rule that fired.
If the gate cannot reach the results, it has to fail loudly rather than default to success. A gate that returns zero when the API is down is worse than no gate at all, because everyone believes it.
How gates fail
Gates rarely fail by being too strict on day one. They fail slowly, by being bypassed, widened, or ignored until they pass everything. Each of these has the same tell: people stop reading the output.
The gate everyone bypasses
If there is an override and it gets used weekly, this is not a gate. Either fix what makes it fire falsely, or remove it and stop pretending.
The gate that only sees automation
Gating on the automated suite while manual results land later means you ship before half the evidence exists. Gate on the run, not on one source of results.
The allowance that only grows
Every time the gate fires, someone raises the threshold by one. Six months later it allows twelve failures and blocks nothing.
The green run nobody executed
A run with 40 unexecuted cases and zero failures satisfies a failure-count rule perfectly. Completion has to be checked on its own.
Signs your gate has stopped working
- The failure allowance has been raised more than once.
- People ask for the bypass by default, not as an exception.
- Nobody can say what rule the gate is currently applying.
- The last three times it fired, the cause was the gate, not the code.
What to measure
A gate is a control, and controls need monitoring of their own. Two numbers tell you whether yours is doing anything useful: how often it fires, and how often it was right to.
- Fire rate
- Share of runs the gate blocks. Near zero means it is decorative. Very high means it is miscalibrated, or the suite is unstable.
- True block rate
- Share of blocks that traced to a real defect. Below about half and people will start bypassing, whatever the policy says.
- Bypass rate
- How often someone overrides the gate. The single best predictor that it is about to be deleted.
- Escaped defects
- Defects found after the gate that the gate could plausibly have caught. This is the number the gate exists to reduce.
Look at these once a quarter. A gate set in March against a suite that has since doubled in size is applying a rule nobody actually chose.
Glossary
- Quality gate
- An automated checkpoint that lets a pipeline continue only when defined conditions are met.
- Exit code
- The number a process returns when it ends. Zero means success. CI systems fail a step on anything else.
- Pass rate
- Passed cases divided by executed cases. Note the denominator: unexecuted cases are not counted, which is why completion needs a separate check.
- Unexecuted case
- A case included in a run that nobody has recorded a result for yet. Neither passed nor failed.
- Flaky test
- A test that passes and fails on the same code. The main reason gates get widened until they stop blocking anything.
- Release gate
- A quality gate placed before a deployment rather than before a merge. Usually allowed to wait on human sign-off.
Further Reading
FAQ
What is a quality gate in software testing?
An automated checkpoint that allows a pipeline to continue only when defined conditions are met. It reads a signal, applies a rule you set, and returns a verdict as an exit code. Non-zero stops the build.
What is the difference between a quality gate and a code coverage threshold?
A coverage threshold is one kind of quality gate, and it reads source code without running it. A test quality gate reads the results of an executed run, so it can tell you whether the software behaved correctly rather than only whether the code was exercised.
Where should a quality gate run in a CI/CD pipeline?
Before merge for fast automated checks, before deploy to staging for longer suites, and before production for anything that needs manual sign-off. The later the gate, the more it can afford to wait.
What should a quality gate do when tests are flaky?
Set a failure allowance and treat it as debt with a date attached, or gate only on completion until the suite is stable. Do not keep raising the allowance. That is how gates stop blocking anything.
Can a quality gate wait for manual testing to finish?
Yes, if it can poll the run. A gate that waits until no case is left unexecuted lets a deploy stage hold for QA sign-off instead of reading a half-finished run and passing it.
What exit code should a quality gate return?
Zero when the gate passes and one when it fails. Use a separate code, commonly two, for usage and API errors, so a gate that could not reach the results is never mistaken for a gate that passed.
Want to gate a build on your test results?
Optional next step: TestCollab reads a test plan's live run and returns an exit code your pipeline already understands, so the same gate works in Azure DevOps, Jenkins, GitLab CI, and GitHub Actions.


