“How many visitors do I need?” is the first question every A/B test should answer — before it starts. Run a test that can never reach significance and you have wasted weeks; the honest answer is that the number depends on three things you control, and you can calculate it in about a minute.
The three inputs that decide the number
- Baseline conversion rate. The current rate of the page you are testing. Rarer events (a 1% baseline) need many more visitors than common ones (a 20% baseline) to measure reliably.
- Minimum detectable effect (MDE). The smallest improvement worth catching. Detecting a jump from 3% to 3.3% (a 10% relative lift) takes far more traffic than catching a jump from 3% to 4.5%.
- Statistical confidence & power. The standard is 95% confidence with 80% power. Demanding more certainty raises the sample size.
Rules of thumb
Exact numbers come from a calculator, but these ballparks help you sanity-check whether a test is even feasible:
- A ~3% baseline, aiming to detect a ~20% relative lift, needs roughly 4,000–5,000 visitors per variant.
- Halve the lift you want to detect and the required sample roughly quadruples.
- Lower baselines (1–2%) can push the requirement into the tens of thousands per variant.
- Two variants split evenly means you need that sample for each — so double it for the whole test.
Turn visitors into a timeline
Once you know the visitors required, divide by your daily traffic to that page to get the duration. If your page gets 300 visitors a day and you need 5,000 per variant (10,000 total), that is about 33 days. Always round up to cover at least one full business cycle (1–2 weeks) so weekday and weekend behavior are both represented.
What to do when you don’t have the traffic
Most pages don’t have unlimited traffic, and that is fine — it just changes what you test:
- Test bigger changes. A new headline, offer, or whole page structure produces a larger effect that needs a smaller sample to prove.
- Test higher up the funnel. Optimize the step with the most volume, not the one with a handful of visitors a day.
- Raise your MDE. Accept that you can only reliably detect meaningful wins, not fractional ones.
- Be patient, not hasty. A longer, honest test beats a fast, false one.
The mistake that wastes the sample: peeking
Deciding the sample size is pointless if you stop the moment a variant crosses 95%. Repeatedly checking and stopping early — “peeking” — inflates false positives badly. Set the sample size and duration in advance, then judge the result once, at the end.
How SplitLab helps
SplitLab tracks unique visitors and conversions per variant and computes statistical significance (a chi-square test) for you, so you always know whether you have enough data to call a winner — no spreadsheets required. Pair it with the significance calculator to plan the test before you launch it.