Skip to content
ScaleFieldLab

Free tool

A/B test significance calculator

An A/B test is significant when z = (rate B − rate A) ÷ standard error gives a p-value below your threshold. 3.00% vs 3.60% on 10,000 visitors each is p = 0.018, significant at 95%.

  • Free
  • No sign-up
  • Runs in your browser

Your test

Results so farVisitors (or sessions, or clicks) and conversions per version.

Rate 3%

Rate 3.6%

Confidence level
Test type
Sample size helperPlan before you start: visitors needed per variant at 80% power, using the confidence and test type above.
Optional
Optional

Both versions combined.

Your results

Relative uplift, B vs A

+20%

3.6% vs 3% (+0.60 pp)

Significant at 95%: B ahead. B converts +20% better (p = 0.018). The true lift is plausibly +0.10 pp to +1.10 pp. You’re at 19% of the planned sample, so don’t stop early.

p-value
0.018
two-sided; needs < 0.050
z-score
2.38
Critical value ±1.96
95% interval for the difference
+0.10 pp to +1.10 pp
Range of plausible true differences
Sample size per variant
53,211
To detect +10% on 3% at 80% power
Time to run
11 weeks
71 days, rounded up to full weeks

How this is calculated
Conversion rate
Conversions ÷ VisitorsWith your numbers: A: 300 ÷ 10,000 = 3% · B: 360 ÷ 10,000 = 3.6%
Relative uplift
(Rate B − Rate A) ÷ Rate AWith your numbers: +0.60 pp ÷ 3% = +20%
Pooled rate
(Conversions A + B) ÷ (Visitors A + B)With your numbers: 660 ÷ 20,000 = 3.3%
z-score
(Rate B − Rate A) ÷ √(Pooled × (1 − Pooled) × (1/Visitors A + 1/Visitors B))With your numbers: +0.60 pp ÷ 0.253 pp = 2.38
p-value
Two-sided: 2 × (1 − Φ(|z|)) · One-sided: 1 − Φ(z)With your numbers: two-sided: 0.018 vs threshold 0.050
Confidence interval
(Rate B − Rate A) ± z* × √(A(1−A)/Visitors A + B(1−B)/Visitors B)With your numbers: 95% two-sided, z* = 1.960: +0.10 pp to +1.10 pp
Sample size per variant
(z_α × √(2p̄(1−p̄)) + z_β × √(p₁(1−p₁) + p₂(1−p₂)))² ÷ (p₂ − p₁)², with 80% powerWith your numbers: Baseline 3%, +10% relative → 53,211 visitors each

Uses the normal approximation to the binomial (a two-proportion z-test), which is reliable once each version has at least 10 or so conversions and non-conversions. Φ is the standard normal distribution; z* is the critical value for your confidence level.

Uplift +20% · p 0.018Results

What does statistical significance mean in plain English?

Two identical pages still convert slightly differently by chance. Statistical significance asks how often a gap this big would appear if A and B were really the same.

Worked example (the calculator's default numbers)
StepValue
A: 300 conversions from 10,000 visitors3.00%
B: 360 conversions from 10,000 visitors3.60%
Relative uplift(3.60 − 3.00) ÷ 3.00 = +20%
Pooled rate and standard error3.30%, 0.253 percentage points
z-score0.60 ÷ 0.253 = 2.38
p-value (two-sided)0.018, significant at 95%
  • The p-value is that probability. Below 0.05 (at 95% confidence), a gap this large would be unusual if the versions performed the same.
  • The confidence interval is often more useful: the range of differences that fit the data. If it runs from +0.1 to +1.1 percentage points, B is probably better, but by an uncertain amount.

How many visitors do you need for an A/B test?

More than most people expect, and more still for low baseline rates and small lifts. Decide the sample size before you start, then run to it.

Visitors needed per variant at 95% confidence (two-sided) and 80% power
Baseline rateDetect +10% relativeDetect +20% relativeDetect +30% relative
1%≈163,100≈42,700≈19,800
3%≈53,200≈13,900≈6,500
5%≈31,200≈8,200≈3,800
10%≈14,800≈3,800≈1,800

So low-traffic UAE sites should test bold changes, such as a new offer, page structure or an Arabic-first landing page, rather than button colours.

What are the most common A/B testing mistakes?

  • Peeking and stopping early. Stopping the moment p dips below 0.05 hugely raises the chance of a false winner. Fix the sample size and end date first.
  • Not running full weeks. UAE traffic shifts on weekends, payday, Ramadan evenings and sale periods. Run whole weeks, ideally two or more.
  • Too many variants or metrics. Test five versions or ten metrics at 95% and something will look significant by luck. Pick one primary metric up front.
  • The novelty effect. Returning visitors click on anything new for a while. A lift that fades after two weeks was novelty.
  • Uneven splits. If a 50/50 test ends up 55/45, check the setup before trusting the result.

Not significant isn't the same as no difference

It means the test couldn't separate the versions with this much data. The true effect could still be positive, negative or zero. If the confidence interval is wide, you need more traffic or a bolder change.

How does this apply to Meta ad creative tests?

The maths works for ads, with impressions or clicks as visitors. But Meta sends more budget to the ad it predicts will win, so in a normal campaign the two ads see different audiences.

  • Use Meta’s A/B test feature (in Experiments) for a clean answer: each person sees only one version.
  • Judge on cost per purchase or lead, not CTR. A creative can win clicks and lose sales.
  • Start with a higher-volume event such as add to cart for early creative rounds, as purchases need far more spend to reach significance.

We run structured creative and landing page testing as part of our performance marketing service.

FAQ

A/B testing questions, answered

Next step

Want tests that give you real answers?

Book a free 30-minute intro call, and we’ll look at your traffic, conversion rates and what is worth testing first.