Free Email ToolsFree Tool

Free Email A/B Test Significance Calculator

Compare two email conversion rates with a free two-proportion significance calculator. See absolute rates, relative uplift, p-value, assumptions, and practical next steps.

Updated August 25, 20268 min readReviewed against linked primary sources
Direct answer

Enter the recipients and conversions for variants A and B. The calculator compares their proportions with a two-sided pooled z-test and reports rates, relative uplift, and an approximate p-value. Use the result only when assignment was random, the outcome and stopping rule were chosen before launch, and each group has enough successes and non-successes for the approximation.

What you will be able to do
  • Compare conversion proportions, not raw conversion counts.
  • Check the large-sample condition before using the z-test result.
  • A p-value is not the probability that the winning variant is better.
  • Predefine one primary outcome and a stopping rule before launch.
  • Pair statistical evidence with the size and business value of the observed effect.
Free browser-based tool

Compare two email test results

Uses a two-sided, pooled two-proportion z-test. It estimates whether the result is distinguishable from random variation; it does not prove causation.

Variant A
Variant B
Variant A rate
5.00%
Variant B rate
6.00%
Relative uplift
20.00%
Two-sided p-value
0.0283
The observed difference is statistically significant at the 95% level.

Use a conversion event tied to the campaign goal, keep assignment random, decide the stopping rule before launch, and do not repeatedly check a running test until it happens to cross 0.05.

A numerical lead is not automatically a reliable win. If variant B converts 6 percent and variant A converts 5 percent, the decision depends on how many recipients were assigned, how the test was run, and whether that one-point difference is valuable enough to act on.

The calculator above uses the large-sample two-proportion method documented by NIST. It is a decision aid for a completed, randomly assigned test, not a substitute for experimental design or a license to keep checking until a favorable result appears.

Choose the conversion before the test starts

A subject-line test often defaults to opens because the event happens near the tested element. That is convenient, but the primary outcome should match the decision. If the campaign exists to create qualified demo requests, a click or booked demo may matter more than an instrumented open.

Choose one primary outcome before launch. Additional metrics can diagnose why a result occurred, but switching the winner criterion after viewing the data makes the conclusion easier to manipulate.

  • Subject or sender test: opens may be directional; clicks can show downstream quality.
  • CTA or body test: unique clicks or qualified conversions usually fit better.
  • Lifecycle test: activation, retention, purchase, or goal exit should match the journey.
  • Revenue test: define the attribution window and eligible order before assignment.

How the calculator compares two rates

Each variant rate is conversions divided by assigned recipients. The calculator pools the two observed proportions under the null hypothesis that the underlying rates are equal, estimates the standard error of their difference, and converts that difference to a z-score and two-sided p-value.

A p-value below 0.05 means a difference at least this large would be relatively unusual under the equal-rate model and the assumptions used. It does not mean there is a 95 percent probability that B is better, and it does not tell you whether the effect is commercially important.

The calculation used
rate A = conversions A / recipients A
rate B = conversions B / recipients B
pooled rate = total conversions / total recipients
z = (rate B - rate A) / pooled standard error
p-value = two-sided probability beyond |z|

Check whether the approximation is appropriate

The large-sample z approximation becomes unstable when success or non-success counts are very small. This page requires at least 10 conversions and 10 non-conversions in each group before it presents the 95-percent verdict as usable.

That threshold is a practical guardrail, not a guarantee. Very rare events, clustered recipients, repeated observations from the same person, unequal assignment mechanisms, or delayed conversions may require a more suitable analysis.

Avoid the early-stopping trap

Repeatedly checking a conventional fixed-horizon p-value and stopping the first time it falls below 0.05 raises the chance of a false positive. Decide the sample, duration, and analysis point before launch, then let the planned window finish unless a safety or operational issue requires stopping.

Also avoid extending only losing tests until they become favorable. If tests must support continuous monitoring, use a sequential method designed for that purpose and document it before looking at the result.

  1. Define the eligible audience and assignment unit.
  2. Select the primary conversion and measurement window.
  3. Estimate the sample needed for the smallest useful effect.
  4. Launch variants concurrently with random assignment.
  5. Analyze at the planned point and record the decision.

Read effect size beside significance

Absolute difference and relative uplift answer different questions. Moving from 5 percent to 6 percent is a one-percentage-point absolute lift and a 20-percent relative lift. Both are true, but the absolute change is often more useful for forecasting incremental conversions.

Translate the observed effect into expected value, cost, risk, and operational complexity. A tiny effect can become statistically significant in a very large audience while remaining too small to justify a permanent workflow change. A promising large effect in a small sample may deserve another planned test rather than an immediate rollout.

Document what the result can and cannot support

Record the hypothesis, variants, audience rules, assignment date, sample, primary outcome, exclusions, result, and decision. This prevents a later report from presenting a directional test as stronger evidence than it was.

If the tested element changed more than one thing, describe the result as a comparison of complete variants. Do not attribute the effect to a single word, layout block, or personalization field that was not isolated.

Frequently asked questions

What does a p-value of 0.05 mean in an email A/B test?

Under the equal-rate model and the test assumptions, a result at least this extreme would occur about 5 percent of the time. It is not a 95 percent probability that the observed winner is truly better.

How many recipients do I need for an email A/B test?

It depends on the baseline conversion rate, smallest useful effect, desired false-positive rate, and statistical power. The calculator checks only whether the completed sample is large enough for its approximation; it does not perform prospective sample-size planning.

Can I use this calculator for open rates?

Mathematically, yes, if each recipient contributes one binary outcome. Interpret opens cautiously because mailbox privacy and image-loading behavior can affect the recorded event. A click or business conversion may better match the campaign goal.

Put the guide into practice

Run the next test with the decision recorded first

Prepare controlled campaign variants, review every recipient email, and keep the result connected to the campaign goal rather than an isolated vanity metric.

21 days, no credit card