Most A/B Tests Are Run Incorrectly
The most common A/B testing mistake is stopping a test early — when one variant appears to be winning — before reaching statistical significance. This "peeking" problem causes teams to ship changes that appear to win in the short run but are actually statistical noise. The result: a CRO programme that doesn't actually improve conversion rates over time.
Statistical Significance: What It Actually Means
Statistical significance at 95% means: if there were no real difference between variants, you would see results this extreme only 5% of the time. It does not mean your test is 95% likely to be correct — it means you've set a bar for how much evidence you require before drawing a conclusion.
A 95% significance threshold with 80% statistical power (the ability to detect a true difference when one exists) requires a specific sample size that must be calculated before the test starts, not after.
Sample Size Calculation
Before any test, calculate the required sample size. Inputs needed:
- Baseline conversion rate: Current conversion rate of the page or step being tested
- Minimum detectable effect (MDE): The smallest improvement worth shipping — typically 10-20% relative improvement
- Significance level: 95% is standard
- Statistical power: 80% is standard
Use a sample size calculator (AB Testguide, Evan Miller's calculator, or your testing tool's built-in calculator). If the required sample size is 10,000 users per variant and your page gets 500 visitors per month, you need to run the test for 40 months — which is not feasible. Focus tests on higher-traffic pages.
Test Duration
Run tests for at least 2 business cycles (typically 2 weeks minimum) regardless of when you reach your sample size. This accounts for weekly seasonality — a test that runs only Monday-Thursday may show different results than a test including the weekend.
What to Do When a Test Loses
A losing test is not a failed test — it's information. A lost test tells you that your hypothesis was wrong, which eliminates a direction from the hypothesis backlog. The learning is as valuable as a win if it updates your model of user behaviour. Document every test result (win, loss, or inconclusive) in a central test log. Our CRO service maintains a test log for every client as part of the ongoing programme.
Need expert tracking setup?
Our Google Tag Manager experts have delivered 500+ tracking setups with a 98% success rate.
Get a Free Consultation →