A/B Test Significance Calculator
Check whether the difference between control and variant is larger than chance would explain.
p-value
0.0768
Enter your numbers and press Calculate.
z-score = (Variant rate − Control rate) ÷ Standard error
Worked example
5,000 / 250 vs 5,000 / 290 p = 0.0768
What this number tells you
The test starts from a null hypothesis: that both sides have the same true conversion rate and the difference you saw is noise from random assignment. The z-test measures how far the observed difference sits from zero, and the p-value says how likely a difference at least that large would be if the null were true.
A p-value of 0.05 does not mean a 95% probability the variant is better. It is a statement about the data given the hypothesis, not about the hypothesis given the data. It also says nothing about the size of the effect.
Statistical significance and practical significance are separate questions. A significant 0.1-point lift may not be worth the complexity of shipping it; a commercially meaningful lift can fail to reach significance simply because the sample was too small.
When to use it
Once, after the test has reached the sample size you planned for. Checking it as data accumulates is what the number cannot survive.
Where it misleads
Repeated peeking and stopping at the first crossing inflates the false-positive rate well past the stated threshold. And not significant does not mean no difference exists; it can equally mean the test was too small to see the one that does.
Frequently asked questions
My test isn't significant. Does that mean there's no difference?
A non-significant result can mean there's no meaningful difference between control and variant, or that the sample was too small to detect the one that exists; the two look identical from the p-value alone. Check the sample size calculator to see whether the traffic and duration this test ran with gave it a real chance of detecting an effect of the size you were hoping to find; if the required sample was much larger than what the test actually collected, 'not significant' is more likely an underpowered test than a true zero effect. A significant result and an underpowered non-significant one aren't symmetric conclusions, even though the calculator's binary significant/not-significant output can make them look that way at a glance.
How long should I run the test before checking this?
Long enough to reach the sample size the test needs for the effect you're trying to detect (the sample size calculator gives that number), and for at least one full weekly cycle, so day-of-week traffic patterns don't skew either side. Stopping as soon as the running p-value first dips below 0.05, rather than waiting for the pre-planned sample size, is one of the most common ways a test result ends up misleading, since checking repeatedly and stopping at the first significant-looking moment inflates the real false-positive rate well above 5%. The test duration estimator converts a required sample size into an actual number of days given current traffic, which is a more reliable stopping rule than watching the p-value move day to day.
What does the p-value actually mean?
The p-value is the probability of seeing a difference at least as large as the one observed, if control and variant actually had the same true conversion rate and the only reason for the observed gap was random chance in who got assigned to each side. It is not the probability that the null hypothesis is true, and it is not the probability that the variant is actually better; both are common misreadings of the same number. A p-value of 0.0768, for example, means that if there were truly no difference between control and variant, a gap this large or larger would show up in roughly 7.7% of random assignments by chance alone. That's a statement about the data given an assumption, not a verdict on which version is actually better.
What's the difference between statistical and practical significance?
Statistical significance asks whether an observed difference is larger than plausible chance variation, given the sample size collected. Practical significance asks whether that difference is large enough to matter for the business, once implementation cost, complexity, and any tradeoffs are considered. The two can disagree in either direction: a tiny, commercially trivial lift can reach statistical significance with a large enough sample, and a genuinely large, valuable lift can fail to reach significance if the sample was too small. A statistically significant result is a signal worth taking seriously, not a shipping decision on its own; the effect size and its confidence interval matter as much as whether the p-value cleared a threshold.
Does 95% confidence mean there's a 95% chance the variant is better?
This is one of the most common misreadings of significance testing. 95% confidence (a p-value below 0.05) means that if control and variant truly had the same conversion rate, a difference this large or larger would appear by chance alone less than 5% of the time. It says nothing directly about the probability that the variant is actually better; that would require a different kind of calculation entirely: a Bayesian approach, which estimates the probability one variant beats another directly rather than testing against a null hypothesis. This frequentist test and a Bayesian one can produce different-sounding answers to what feels like the same question, because they're formally answering different questions.
Related calculators