A numerical lead is not automatically a reliable win. If variant B converts 6 percent and variant A converts 5 percent, the decision depends on how many recipients were assigned, how the test was run, and whether that one-point difference is valuable enough to act on.
The calculator above uses the large-sample two-proportion method documented by NIST. It is a decision aid for a completed, randomly assigned test, not a substitute for experimental design or a license to keep checking until a favorable result appears.
Choose the conversion before the test starts
A subject-line test often defaults to opens because the event happens near the tested element. That is convenient, but the primary outcome should match the decision. If the campaign exists to create qualified demo requests, a click or booked demo may matter more than an instrumented open.
Choose one primary outcome before launch. Additional metrics can diagnose why a result occurred, but switching the winner criterion after viewing the data makes the conclusion easier to manipulate.
- Subject or sender test: opens may be directional; clicks can show downstream quality.
- CTA or body test: unique clicks or qualified conversions usually fit better.
- Lifecycle test: activation, retention, purchase, or goal exit should match the journey.
- Revenue test: define the attribution window and eligible order before assignment.
How the calculator compares two rates
Each variant rate is conversions divided by assigned recipients. The calculator pools the two observed proportions under the null hypothesis that the underlying rates are equal, estimates the standard error of their difference, and converts that difference to a z-score and two-sided p-value.
A p-value below 0.05 means a difference at least this large would be relatively unusual under the equal-rate model and the assumptions used. It does not mean there is a 95 percent probability that B is better, and it does not tell you whether the effect is commercially important.
rate A = conversions A / recipients A
rate B = conversions B / recipients B
pooled rate = total conversions / total recipients
z = (rate B - rate A) / pooled standard error
p-value = two-sided probability beyond |z|Check whether the approximation is appropriate
The large-sample z approximation becomes unstable when success or non-success counts are very small. This page requires at least 10 conversions and 10 non-conversions in each group before it presents the 95-percent verdict as usable.
That threshold is a practical guardrail, not a guarantee. Very rare events, clustered recipients, repeated observations from the same person, unequal assignment mechanisms, or delayed conversions may require a more suitable analysis.
Avoid the early-stopping trap
Repeatedly checking a conventional fixed-horizon p-value and stopping the first time it falls below 0.05 raises the chance of a false positive. Decide the sample, duration, and analysis point before launch, then let the planned window finish unless a safety or operational issue requires stopping.
Also avoid extending only losing tests until they become favorable. If tests must support continuous monitoring, use a sequential method designed for that purpose and document it before looking at the result.
- Define the eligible audience and assignment unit.
- Select the primary conversion and measurement window.
- Estimate the sample needed for the smallest useful effect.
- Launch variants concurrently with random assignment.
- Analyze at the planned point and record the decision.
Read effect size beside significance
Absolute difference and relative uplift answer different questions. Moving from 5 percent to 6 percent is a one-percentage-point absolute lift and a 20-percent relative lift. Both are true, but the absolute change is often more useful for forecasting incremental conversions.
Translate the observed effect into expected value, cost, risk, and operational complexity. A tiny effect can become statistically significant in a very large audience while remaining too small to justify a permanent workflow change. A promising large effect in a small sample may deserve another planned test rather than an immediate rollout.
Document what the result can and cannot support
Record the hypothesis, variants, audience rules, assignment date, sample, primary outcome, exclusions, result, and decision. This prevents a later report from presenting a directional test as stronger evidence than it was.
If the tested element changed more than one thing, describe the result as a comparison of complete variants. Do not attribute the effect to a single word, layout block, or personalization field that was not isolated.