Free tool

A/B test sample-size & duration calculator

Set your baseline and the minimum detectable effect (MDE) to work out how many visitors - and how many days - you'll need before a test can give a trustworthy answer. Plan the sample up front and you won't stop too early, the single most common A/B testing mistake.

Power analysis
%
%
Your plan

-

Enter your baseline rate and the lift you want to detect.

-visitors per variant
-visitors in total, two variants
-days at your traffic

How this works

The calculator uses a standard power analysis for comparing two proportions, at 80% statistical power. It is the same method as the planner inside AB Test WP: a two-sided test at 95% confidence and 80% power, the lift expressed relative to your baseline, the total being the per-variant figure times the number of variants, and the days being that total divided by your daily visitors. In plain terms: it finds the smallest sample where, if a real improvement of the size you chose exists, you'd reliably detect it without being fooled by random noise.

  • Smaller effects need far more traffic. Detecting a 5% lift takes many times more visitors than a 50% lift.
  • Run for whole weeks. Even once you hit the sample size, let the test cover full weekly cycles so weekday/weekend behaviour evens out.
  • Decide the sample before you start, then wait for it - don't stop the moment a variant looks ahead.
When your test is live, check it with the significance calculator. AB Test WP includes this planner inside the plugin, so you can size a test without leaving WordPress.

FAQ

Questions

What do confidence and power mean here?

Confidence (95% by default) controls false positives: how sure you want to be before calling a difference real. Power (80% here) controls false negatives: the chance the test detects the uplift if it truly exists. Together they set the sample size.

Is the uplift relative or absolute?

Relative. A 20% uplift on a 3% baseline means detecting a move from 3% to 3.6% - not from 3% to 23%. Smaller relative uplifts need dramatically more traffic to detect.

What if my site can't reach the sample size?

Test bolder changes. Big differences (a new headline and offer, a rebuilt page) are much easier to detect than small tweaks, so low-traffic sites should run fewer, braver tests - and be comfortable calling close results inconclusive.

Quick reference: 3% baseline, 95% confidence, 80% power

Visitors needed per variant at a 3% baseline conversion rate
Relative lift to detectPer variantTotal (A/B)
+5% (3% → 3.15%)≈ 208,000≈ 416,000
+10% (3% → 3.3%)≈ 53,000≈ 106,000
+20% (3% → 3.6%)≈ 14,000≈ 28,000
+50% (3% → 4.5%)≈ 2,500≈ 5,000

Notice the shape: halving the effect you want to detect roughly quadruples the traffic you need. This is why low-traffic sites should test bold changes rather than small tweaks.

Related reading: how long to run an A/B test and A/B testing statistics made simple.

Plan it, then test it

The free plugin ships with this planner built in - install it and run your first proper test.