A/B Testing Fundamentals: Sample Size and Significance | Adslytics | Adslytics

Behavior Analytics & CRO Educational

A/B Testing Fundamentals: Sample Size, Statistical Significance, and Duration

By Muhammad Farooq · July 14, 2026 · 8 min read
A/B Testing Fundamentals: Sample Size, Statistical Significance, and Duration

Why Most A/B Tests Are Invalid

The majority of A/B tests in the wild are statistically invalid. They're stopped too early, run on too little traffic, or declared winners based on metrics that weren't defined before the test started. The result is "winning" tests that fail when implemented, wasted development time, and false confidence in changes that didn't actually work.

Statistical validity is not complicated, but it requires discipline. Our CRO team follows these fundamentals on every test we design.

The Three Statistical Concepts You Need

Statistical Significance (p-value)

Statistical significance tells you how likely it is that your results happened by chance. A p-value of 0.05 (95% significance) means: if there were actually no difference between control and variant, you'd see results this extreme only 5% of the time by random chance.

The common misconception: "95% significant" does NOT mean "95% certain the variant is better." It means "5% probability these results are random noise given no true effect." That's a meaningful distinction when you run many tests.

Standard threshold: 95% significance (p < 0.05) is the minimum for most CRO tests. Use 99% for high-stakes decisions (major design overhauls, pricing changes).

Statistical Power

Power is the probability of detecting a real effect if one exists. Low-powered tests miss real improvements — they fail to reject the null hypothesis even when the variant is genuinely better.

Standard power: 80% (meaning 20% chance of a false negative — missing a real winner). Higher power requires more sample size.

Minimum Detectable Effect (MDE)

The smallest improvement you care about detecting. If your current conversion rate is 3% and you'd only implement a change that improves it to 3.5% or more, your MDE is 0.5 percentage points (or ~17% relative improvement).

Smaller MDE requires larger sample size. This is the most common mistake in CRO: testing to detect a 5% lift when you'd need 100,000 visitors per variation to do it reliably.

Calculating Required Sample Size

Use a sample size calculator (Evan's Awesome A/B Tools, AB Testguide, or any statistical power calculator). Inputs:

  • Baseline conversion rate: your current rate on the control page
  • Minimum detectable effect: the smallest lift you care about
  • Statistical significance: typically 95%
  • Statistical power: typically 80%

Example: Baseline rate 3%, MDE 1 percentage point (from 3% to 4%), 95% sig, 80% power → ~4,600 visitors per variation (9,200 total)

At 500 daily visitors per variation, this test needs ~18 days to run.

Setting Test Duration

Test duration is not just "run until significant." It's determined by two factors:

  1. Statistical requirement: Time needed to collect required sample size
  2. Business cycles: Minimum 1–2 full weeks to account for day-of-week variation (Monday behavior differs from Friday behavior)

Run for whichever is longer. A test that reaches significance after 3 days should still run for a full week minimum to avoid "day-of-week bias."

Maximum duration: 4–6 weeks. After that, external factors (seasonality, news, campaigns) start contaminating results.

The Peeking Problem

The most common mistake: checking results daily and stopping the test as soon as significance is reached. This is called "peeking" and dramatically inflates your false positive rate.

In a simulation: if you check significance at every day of a 4-week test and stop when you first hit 95%, your actual false positive rate is ~30% — not 5%. You're 6x more likely to declare a false winner than the p-value suggests.

Solution options:

  • Set the end date before starting. Don't look at results until the test completes.
  • Use Sequential Testing / Bayesian methods that are designed for early stopping
  • Apply Bonferroni correction if you must check at multiple points

One Test, One Primary Metric

Define one primary success metric before the test starts. Secondary metrics can be observed, but the test is only declared a winner/loser based on the primary metric.

Testing against multiple metrics without correction inflates your false discovery rate significantly. If you test 10 metrics at p < 0.05, you'd expect 0.5 false positives by chance — in a single 10-metric analysis, one of those "significant" results is likely noise.

Connect your A/B testing tool to GA4 for conversion event tracking. Our CRO team designs statistically rigorous experiments. Contact us to set up a proper testing program.

Need expert tracking setup?

Our Google Tag Manager experts have delivered 500+ tracking setups with a 98% success rate.

Get a Free Consultation →
← Back to Blog
Muhammad Farooq

Author

Muhammad Farooq GTM & Analytics Expert · Adslytics Founder

Tracking specialist with 10+ years of experience in Google Tag Manager, GA4, Server-Side Tracking, and Google Ads. Founder of Adslytics — a dedicated analytics agency with a 98% success rate across 232+ projects on Upwork.

Top Rated Plus LinkedIn Visit the author's profile →