Free tool
A/B test sample-size & duration calculator
Set your baseline and the minimum detectable effect (MDE) to work out how many visitors - and how many days - you'll need before a test can give a trustworthy answer. Plan the sample up front and you won't stop too early, the single most common A/B testing mistake.
-
Enter your baseline rate and the lift you want to detect.
Your inputs are saved in the page address.
How this works
The calculator uses a standard power analysis for comparing two proportions, at 80% statistical power. It is the same method as the planner inside AB Test WP: a two-sided test at 95% confidence and 80% power, the lift expressed relative to your baseline, the total being the per-variant figure times the number of variants, and the days being that total divided by your daily visitors. In plain terms: it finds the smallest sample where, if a real improvement of the size you chose exists, you'd reliably detect it without being fooled by random noise.
- Smaller effects need far more traffic. Detecting a 5% lift takes many times more visitors than a 50% lift.
- Run for whole weeks. Even once you hit the sample size, let the test cover full weekly cycles so weekday/weekend behaviour evens out.
- Decide the sample before you start, then wait for it - don't stop the moment a variant looks ahead.
FAQ
Questions
What do confidence and power mean here?
Confidence (95% by default) controls false positives: how sure you want to be before calling a difference real. Power (80% here) controls false negatives: the chance the test detects the uplift if it truly exists. Together they set the sample size.
Is the uplift relative or absolute?
Relative. A 20% uplift on a 3% baseline means detecting a move from 3% to 3.6% - not from 3% to 23%. Smaller relative uplifts need dramatically more traffic to detect.
What if my site can't reach the sample size?
Test bolder changes. Big differences (a new headline and offer, a rebuilt page) are much easier to detect than small tweaks, so low-traffic sites should run fewer, braver tests - and be comfortable calling close results inconclusive.
Quick reference: 3% baseline, 95% confidence, 80% power
| Relative lift to detect | Per variant | Total (A/B) |
|---|---|---|
| +5% (3% → 3.15%) | ≈ 208,000 | ≈ 416,000 |
| +10% (3% → 3.3%) | ≈ 53,000 | ≈ 106,000 |
| +20% (3% → 3.6%) | ≈ 14,000 | ≈ 28,000 |
| +50% (3% → 4.5%) | ≈ 2,500 | ≈ 5,000 |
Notice the shape: halving the effect you want to detect roughly quadruples the traffic you need. This is why low-traffic sites should test bold changes rather than small tweaks.
Related reading: how long to run an A/B test and A/B testing statistics made simple.
Plan it, then test it
The free plugin ships with this planner built in - install it and run your first proper test.