LanderKit

Templates written in French — fully translatable in minutes

Statistical significance in a landing page A/B test: what threshold should you use?

Published on 12 August 2026 · 8 min read

An A/B testing tool shows "95% confidence, variant B wins," and the temptation is immediate: roll the change out across the whole page. The number feels reassuring — almost a guarantee. It isn't quite, for two reasons most tools never explain: what that percentage actually measures, and what happens when you've looked at several metrics or segments, maybe without even noticing, before it showed up. Our article on how long to run an A/B test covers when to read the result; this one covers what the result actually means once you read it — and the threshold to require before trusting it.

What a confidence percentage actually measures

"95% confidence" does not mean "95% chance variant B is genuinely better." It means: if the two variants actually converted at exactly the same rate, there would only be a 5% chance of seeing by pure luck a gap as large as the one you measured. It's a measure of noise, not a measure of certainty about the cause. That distinction sounds technical, but it has a concrete consequence: the more opportunities you give randomness to produce a flattering result, the more of those you'll see — even when nothing has actually changed on your page.

The real risk: stacking up chances to be wrong

A test comparing a single metric between two variants carries roughly a 5% false-positive risk at the 95% threshold. But few marketers stop at one comparison: you check the overall conversion rate, then mobile, then desktop, then visitors from Google Ads, then the CTA click rate on top of the final conversion rate. Each of those readings is a fresh opportunity for randomness to produce something that looks significant. The overall risk of seeing at least one false "winner" climbs fast, well past the 5% shown on the single reading that convinced you.

This phenomenon — known as researcher degrees of freedom, the freedom to decide after the fact which metric or subgroup to look at — is documented in detail by a study by Joseph Simmons, Leif Nelson and Uri Simonsohn, published in 2011 in Psychological Science. Their simulations show that combining several flexible analysis choices, each innocuous in isolation — adding a few observations, trying two or three ways of splitting a subgroup, testing an alternative metric — pushes the real false-positive rate well above the nominal 5% shown by the tools. In other words: each individual reading playing by the rules doesn't guarantee your whole set of readings stays reliable.

Multiple comparisons have a cost — and a correction

This isn't a marketing-specific problem: statisticians have known it for decades as the multiple comparisons problem. The fix isn't to ignore the risk, it's to raise the bar as you stack up readings. A landmark study by Yoav Benjamini and Yosef Hochberg, published in 1995 in the Journal of the Royal Statistical Society, proposes a method still used today to control the rate of false discoveries when testing several hypotheses at once: instead of requiring a flat 5% threshold for each comparison taken in isolation, you adjust the required threshold based on how many comparisons were made. You don't need to implement the formula yourself — the takeaway is simpler: if you're checking several metrics or segments within the same test, require a clearer gap on any single one of them than you would on the overall result before believing it.

95% confidence isn't the same thing as a business win

A result can be statistically significant — hard to explain by chance alone — without being worth shipping. With enough traffic, a 0.2-point gap in conversion rate can clear any significance threshold while representing a negligible gain once weighed against revenue impact or the development time needed to roll it out everywhere. Conversely, a test stopped too early can miss a real gain simply for lack of sample size — see our article on the sample size a reliable test needs. The significance threshold answers "is this result noise?" It does not answer "is this result worth deploying?" — that question gets settled by pricing out the expected gain, not by reading a percentage.

Which threshold to use, based on what's at stake

There is no single universal "correct" threshold — the right level depends on what a mistake costs you. Three practical benchmarks for a small team without an in-house statistician or massive traffic volumes.

Choosing your confidence threshold based on the stakes of the change being tested
ThresholdWhen to use itWhat you're accepting
90%A reversible, low-cost change (button color, CTA wording) you can iterate on quicklyA few more false positives, offset by faster iteration
95%The default threshold for most landing page decisions: page structure, featured offer, section orderThe standard balance between rigor and decision speed
99%A costly or hard-to-undo change: checkout flow redesign, displayed price change, removing a form stepA near-certain signal required before acting — you wait longer, but avoid rolling out a mirage

The practical rules, in order

  1. Set the confidence threshold before launching the test, based on the stakes of the change (see the table above) — never after the fact, based on what "looks significant."
  2. Pick one single primary metric to decide the test; other metrics stay informative, not decisive.
  3. If you're looking at segments (mobile vs. desktop, traffic source), require a clearer gap on a single segment than on the overall result before drawing a conclusion from it.
  4. Never confuse a significant result with a result worth shipping: price out the business gain before rolling a change out everywhere.
  5. Stick to the duration and sample size fixed in advance — see our article on A/B test duration — rather than stopping at the first flattering result.
  6. When a result is surprising, re-verify it before believing it, the way large-scale testing programs do.

Rigor isn't a luxury reserved for big companies

A misunderstood confidence threshold creates a false sense of certainty — and nothing is more costly than a landing page decision made on a statistical mirage, then rolled out across your whole acquisition. Set the threshold before launching, limit the number of comparisons that actually drive the decision, and keep statistical significance separate from business relevance. If you're still short on a method for deciding what to test first, our guide on how to prioritize your A/B tests complements this one. And to avoid testing flaws that are already well known, the 10 LanderKit templates ($89 each, $229 for the full bundle) start from a proven structure — short form, well-placed social proof, visible promise — so your tests focus on real hypotheses instead of obvious fixes.

FAQ

Frequently asked questions

What does "95% confidence" actually mean in an A/B test?

It means that if the two variants genuinely converted at the same rate, there would only be a 5% chance of seeing by luck a gap as large as the one measured. It is not "95% chance variant B is better" — it's a measure of noise, not a guarantee of causation.

Why does checking multiple segments or metrics increase the risk of being wrong?

Each additional metric or segment you check (mobile, desktop, traffic source...) is another chance for randomness to produce something that looks significant. A 2011 study by Simmons, Nelson and Simonsohn shows that combining several flexible analysis choices pushes the real false-positive rate well past the 5% shown on any single reading.

What confidence threshold should I use for a landing page A/B test?

95% is a solid default for most decisions. A lower threshold (90%) is acceptable for a reversible, low-cost change you can iterate on quickly. A stricter threshold (99%) is warranted for a costly or hard-to-undo change, like a checkout flow redesign.

Is a statistically significant result always worth shipping?

No. With enough traffic, a tiny gap can clear any significance threshold without representing a meaningful business gain. The significance threshold answers "is this result noise?", not "is it worth deploying?" — that second question is settled by pricing out the expected gain.

Read next

Related articles