The statistics behind this calculator
What it calculates
Each rate is conversions divided by impressions. Absolute lift is the rate of B minus the rate of A, in percentage points. Relative lift is that difference divided by the rate of A, and it is left blank when A has no conversions. The significance test is a two-tailed, pooled two-proportion z-test, with the normal cumulative distribution computed from a high-precision complementary error function. The 95% interval for the difference uses the unpooled standard error and 1.96.
The sample size uses the standard two-proportion formula with a two-tailed 5% significance level and 80% power, evaluated at the rates you entered. It is a planning figure: real rates drift as more data arrives, so recompute as you go.
The Bayesian estimate
Each variant gets a Beta distribution with parameters one plus its conversions and one plus its non-conversions, which is a uniform prior. The tool draws 10,000 samples from each with a fixed random seed, so the same inputs always return the same number, and reports the share of draws where B is higher. It is labeled an estimate because it comes from sampling and depends on the flat prior.
Limitations
The z-test assumes independent impressions and reasonably large counts. It is less reliable with fewer than about 10 conversions in either variant. Platforms may show your variants to different audiences or at different times, which breaks the fair comparison the test assumes. A significant result tells you the gap is unlikely to be chance, not that the cause is the hook or thumbnail you changed. Run one change at a time.
Frequently asked questions
How many views do I need for an A/B test?
It depends on how small a difference you want to detect. With a 5% baseline and a 7% variant, a standard two-proportion calculation at 95% confidence and 80% power needs roughly 2,200 impressions per variant. A smaller gap needs far more, because the required sample grows with the inverse square of the difference. Enter your numbers and the tool shows the figure for your rates.
What does the p-value mean here?
It is the probability of seeing a gap at least this large between A and B if the two variants really converted at the same rate. A p-value below 0.05 is the usual bar for calling a result significant at 95% confidence. It is not the probability that B is better, which is why a separate estimate for that is shown below it.
What is the probability that B beats A?
It is a Bayesian estimate. The tool draws 10,000 seeded random samples from a Beta distribution for each variant, starting from a flat prior, and counts how often B comes out higher. It is a Monte Carlo estimate, so it can differ from an exact figure in the second decimal place, and it answers a different question from the p-value.
Can I stop the test as soon as it turns significant?
It is risky. Checking repeatedly and stopping at the first p-value under 0.05 inflates false positives well above 5%. Decide on a sample size or a duration first, using the sample size shown, and read the result once it is reached.
What counts as a conversion for social posts?
Anything you can count per impression: link clicks, profile visits, saves, replies, signups. Pick one before you start. Comparing likes on one variant with clicks on the other is not a valid test, and the two variants should run at similar times to the same audience.
Why is my result not significant when B looks much better?
With few impressions, a large-looking gap is easily produced by chance. For example, 50 conversions in 1,000 impressions against 70 in 1,000 is a 40% relative lift but gives a p-value near 0.06. The tool shows the interval for the difference, and if it includes zero the data does not rule out no difference at all.