Landing Page Testing

Landing page A/B test sample size and significance: how long to run a test

By the ROAS365 team 2026-07-27 10 min read

Most landing page A/B tests never produce a usable answer. The problem is rarely the creative — it is the measurement: the sample was too small, the test was called the moment it looked good, or invalid traffic polluted the denominators. This guide covers only the measurement layer: the three numbers to fix before you start, what significance actually means, when it is safe to stop, and what quietly invalidates a result.

TL;DR
  • Fix three numbers before you start: your current conversion rate (baseline), the smallest lift worth acting on (the MDE), and daily traffic. Those three determine how long the test runs — not the other way round.
  • The lower the baseline and the smaller the lift you want to detect, the faster the required sample grows. Detecting a 20% relative lift needs roughly 39,000 visitors per variant at a 2% baseline, but only about 6,500 at a 10% baseline.
  • 95% significance does not mean "there is a 95% chance the variant is better". It means that if the two pages were truly identical, a gap this large would appear less than 5% of the time.
  • Run whole weeks until you hit the pre-calculated sample, then read the result. Peeking daily and stopping at the first significant reading is the single biggest source of false winners.
  • Before you conclude anything, rule out invalid traffic. If bots and invalid clicks land unevenly across variants, no significance figure means much.

Why most landing page tests never reach a conclusion

A typical failure looks like this: two versions go live on a Tuesday, by Thursday morning the dashboard shows variant B converting 18% higher at 96% significance, the team ships B, and a week later overall conversion has not moved. The creative was not the issue. The conclusion was never valid — it rested on a few hundred visitors and an early stop.

Landing page tests have three structural difficulties: conversion rates are usually low (1%–5%), real improvements are usually small (5%–20% in relative terms), and paid traffic carries noise of its own (bots, misclicks, cross-device journeys). Stacked together, they mean a visible-looking gap is often just random variation. Without measurement discipline, even a genuinely better page cannot be proven better.

If you have not yet decided what to test or in what order, start with our complete landing page A/B testing guide. This article deals only with getting the answer right.

The three numbers to fix before you start

Do not launch first and ask "how long?" afterwards. Reverse the order: fix three numbers, derive the required sample and duration, then decide whether the test is worth running at all.

1. Baseline conversion rate

Take the real conversion rate of the current page over the last two to four weeks, not the best week you ever had. Write the definition down: is the denominator landing page visitors or ad clicks? Is the numerator a form submit, an add-to-cart or a completed order? An unstable definition distorts every number that follows.

2. Minimum detectable effect (MDE)

Ask a blunt question: how much lift would actually make you change the page and absorb the cost of changing it? If a 5% relative lift would not move you, do not set the MDE at 5% — set it at the number you genuinely care about. The smaller the MDE, the more traffic you need, and that is where most of the measurement cost comes from.

3. Available daily traffic

Use the traffic that will actually enter the test, not total site traffic. If only a few campaigns, only mobile, or only one region is in scope, the denominator shrinks accordingly. Divide the required sample by daily traffic to get the number of days, then round up to whole weeks.

A sample size reference table

The table below gives the approximate number of visitors needed per variant at 95% confidence and 80% power, two-sided. Use it for order-of-magnitude planning, and run your own numbers through a sample size calculator before you commit.

Baseline CR Detect +10% relative Detect +20% relative Detect +50% relative
1%~310,000~81,000~14,000
2%~150,000~39,000~7,000
5%~59,000~15,000~2,700
10%~28,000~6,500~1,300

Approximate visitors required per variant (95% confidence, 80% power, two-sided). Each variant needs this much, so total traffic is double.

The most useful thing this table does is tell you when a test is simply not runnable. A 1% baseline with 500 visitors a day and a 10% MDE needs 310,000 visitors per variant — over three years. The right response is not to run it anyway, but to change the lever: test a bigger change (a whole page rather than button copy), move the conversion event earlier in the funnel (add-to-cart instead of purchase), or grow traffic first.

Common misread

"We get 50,000 visitors" usually refers to total traffic. Each variant receives half of it, and often less once you segment by device, region or campaign. Always size a test on what a single variant will actually receive.

What statistical significance actually says

A 95% significance level (p < 0.05) means precisely this: assuming the two pages are truly identical, a gap at least as large as the one observed would occur less than 5% of the time. It does not mean "there is a 95% chance B is better", and it says nothing about the size of the improvement.

The confidence interval is more useful. "B is 12% higher than A (95% CI: +2% to +23%)" tells you far more than the word "significant" — direction, magnitude and uncertainty in one line. If the interval crosses zero, you have not even established direction.

One quantity gets ignored too often: power. 80% power means that if the true lift equals your MDE, you have an 80% chance of detecting it — and a 20% chance of missing a genuinely better page. An undersized test does not only produce false positives; it systematically buries real winners.

When to stop a test

The rule is simple and the hardest part to follow: decide the sample size and the end date before launch, and read the conclusion on that date. Looking mid-flight is fine; treating a significant reading as permission to stop early is not.

What quietly invalidates a result

Even with the right sample and duration, the following will invalidate a conclusion — and none of them throw an error. They just hand you a wrong answer quietly.

Sample ratio mismatch

If you configured a 50/50 split but the data shows 52/48, something is wrong upstream: one variant loads more slowly, a device class is being routed unevenly, or tracking is dropping part of one side. When the ratio deviates noticeably, fix the split before reading anything.

Invalid traffic and bots

Bots almost never convert but do inflate denominators. If they land unevenly across variants, what you measured is the bot distribution, not the page. Screen for it before concluding — see our notes on detecting bot traffic and what counts as invalid traffic.

Redirect-based splits and speed skew

Splitting traffic by redirecting part of it to a second URL adds a hop of latency to that side, which is especially visible on mobile. You think you are testing copy; you are partly testing load time. Serving both versions within one URL avoids that bias — see our same-URL A/B testing methodology.

Changing other variables mid-test

Adjusting budgets, swapping creatives, changing audiences or adding campaigns mid-test changes the composition of traffic entering the test. When the traffic changes, comparability breaks. Either freeze the period or log every change and analyse the segments separately.

Cross-device and delayed conversions

A user who sees variant B on a phone and buys on a desktop often will not be credited to B. The longer the sales cycle, the larger this bias. Leave a lag window for conversions to land before reading the result.

A pre-launch checklist

Walk through this before launch and most wasted tests disappear:

Once a test ends, recording the learning matters more than winning it. Which directions worked, which did not, and by how much — that log decides what you test next. For the levers outside testing itself, see improving landing page conversion rate and how to improve ROAS.

Frequently asked questions

How long should a landing page A/B test run?

Run until you reach the sample size you calculated before starting, and always in whole weeks — a minimum of one full week, more commonly two to four. Whole weeks matter because weekday and weekend traffic convert differently; stopping mid-week lets the day mix, rather than the page, decide the winner.

What sample size do I need for a landing page A/B test?

It depends on your baseline conversion rate and the smallest lift worth detecting. As a rough guide at 95% confidence and 80% power: a 2% baseline needs roughly 39,000 visitors per variant to detect a 20% relative lift, while a 10% baseline needs about 6,500 per variant for the same relative lift. Smaller lifts require dramatically more traffic.

What does 95% statistical significance actually mean?

It means that if the two pages were truly identical, you would see a difference this large or larger less than 5% of the time by chance alone. It is not the probability that your variant is better, and it says nothing about how large the improvement is — read the confidence interval for that.

Why do A/B test winners often fail to hold up after launch?

The most common causes are stopping the test the moment it looked significant, running it for less than a full week, splitting traffic unevenly, or letting invalid and bot traffic land in one variant more than the other. Each of these produces a result that looks like a winner but does not reproduce once the page is live.

Want to test landing pages within a single URL?

Every visitor hits the same landing-page URL — in-page A/B testing, audience-aware content and invalid-traffic filtering, with every served result inspectable in the dashboard. No hidden content, no sneaky redirects.