The Grandmother Test

Do we have enough traffic to A/B test?

Usually less than you hope. To detect a 10% relative lift on a 3% conversion rate at conventional thresholds you need roughly 25,000 visitors per variant — so a site doing 5,000 visitors a month is looking at most of a year for one test. Below that, testing produces confident noise.

Gergana Tyaneva · 21 September 2026 · 12 years in product and marketing analytics

The grandmother version

You want to know whether a coin is fair. You toss it ten times and get six heads.

That tells you almost nothing. A perfectly fair coin does that all the time.

Toss it a thousand times and six hundred heads means something. The problem was never the coin — it was that you didn't toss it enough.

Do the arithmetic first

Sample size is set by three things: your current conversion rate, the smallest lift worth detecting, and how much risk of being wrong you will accept. Lower conversion rates and smaller effects both need dramatically more traffic.

Work this out before designing anything. If the number of weeks required is longer than the season you are testing in, the test cannot answer the question, and finding that out costs nothing.

Visitors needed per variant to detect a 10% relative lift1% conversion78,0002% conversion39,0003% conversion25,0005% conversion15,00010% conversion7,000A site with 5,000 visitors a month and a 3% rate needs about 10 months for one test.Figures are order-of-magnitude anchors at conventional thresholds, not a substitute for a calculator.

What to do with too little traffic

Test bigger things. A radical redesign has a far larger effect size than a button colour, and large effects need far less traffic to detect.

Move up the funnel, where volumes are higher. Or stop testing and do qualitative work — five user interviews will teach a low-traffic site more than an underpowered test, and won't produce a false certainty.

Call a flat result flat

The discipline that separates a real testing programme from theatre is being willing to say "no detectable difference" and mean it.

Peeking at results until they cross significance, then stopping, guarantees you will find effects that are not there. Fix the sample size and the end date before launch, and honour them.

The short version

Underpowered tests do not give you a weak answer. They give you a confident wrong one, roughly as often as a coin.

If the arithmetic says you cannot detect it, that is useful information delivered before you spent six weeks.

Not sure whether your data can answer the question at all? The Data Readiness Check is €400.

See prices →

More answers

Attribution — Which channel actually pays?Plumbing — Our numbers live in eight tools and no two agreeTime — Month-end reporting eats three days and nobody reads itMoney — We know people churn. We don't know who, when, or what it's worth

One question before you read on. I'd like to count visits with Google Analytics, which sets cookies and sends your IP to Google. Nothing has loaded yet, and the site works exactly the same either way. What's collected.