Did I reach statistical significance?
Type the visitors and conversions of version A and version B. The statistical significance calculator compares the two conversion rates with a two-proportion z-test and gives the z statistic, the p-value, the uplift and whether the difference is significant at 90%, 95% or 99% confidence.
- Result
- Significant at 95%: B converts better
With 200 of 10,000 against 250 of 10,000, p = 0.01713: Significant at 95%: B converts better.
- p-value
- 0.01713
- z statistic
- 2.384
- Conversion rate A
- 2%
- Conversion rate B
- 2.5%
- Relative uplift
- 25%
- Difference (points)
- 0.5
- Pooled rate
- 2.25%
- Significance level α
- 0.05
Result: Significant at 95%: B converts better. With 200 of 10,000 against 250 of 10,000, p = 0.01713: Significant at 95%: B converts better.
Where does your z fall?
How to calculate
Tests an A/B result for statistical significance: two conversion rates compared with the two-proportion z-test, giving the z statistic, the p-value and whether the difference is significant at 90%, 95% or 99%.
Example with the default inputs (Visitors in A 10,000, Conversions in A 200, Visitors in B 10,000, Conversions in B 250, Confidence 95%, Test Two-sided (A ≠ B)): With 200 of 10,000 against 250 of 10,000, p = 0.01713: Significant at 95%: B converts better.
Method: p̂ = (x_A + x_B) ÷ (n_A + n_B); z = (x_B/n_B − x_A/n_A) ÷ √(p̂(1 − p̂)(1/n_A + 1/n_B)); p = 2Φ(−|z|) two-sided or Φ(−z) one-sided; significant when p < 1 − confidence.
- A large-sample test: the normal approximation holds when each group has several conversions and several non-conversions.
- Visitors are independent and each is counted in one group only.
- The one-sided test asks only whether B is better than A.
Worked examples
Each example is checked against the calculator on every build.
- Visitors in A 38, Conversions in A 32, Visitors in B 44, Conversions in B 39, Confidence 95%, Test Two-sided (A ≠ B) gives z statistic 0.586401, p-value 0.557606, Result Not significant at 95%.Source: NIST Dataplot Reference Manual, Binomial Proportion Test (32 of 38 against 39 of 44: pooled 0.86585, statistic −0.58640, two-tailed p-value 0.55760), https://itl.nist.gov/div898/software/dataplot/refman1/auxillar/binotest.htm (retrieved 2026-10-02); NIST/SEMATECH e-Handbook of Statistical Methods, §7.3.3 How can we determine whether two processes produce the same proportion of defectives? (large samples: z = (p̂₁ − p̂₂) ÷ √(p̂(1 − p̂)(1/n₁ + 1/n₂)), p̂ = (x₁ + x₂) ÷ (n₁ + n₂)), https://www.itl.nist.gov/div898/handbook/prc/section3/prc33.htm (retrieved 2026-10-02)
- Visitors in A 10,000, Conversions in A 200, Visitors in B 10,000, Conversions in B 250, Confidence 95%, Test Two-sided (A ≠ B) gives z statistic 2.383995, p-value 0.017126, Relative uplift 25%, Result Significant at 95%: B converts better.Source: NIST/SEMATECH e-Handbook of Statistical Methods, §7.3.3 How can we determine whether two processes produce the same proportion of defectives? (large samples: z = (p̂₁ − p̂₂) ÷ √(p̂(1 − p̂)(1/n₁ + 1/n₂)), p̂ = (x₁ + x₂) ÷ (n₁ + n₂)), https://www.itl.nist.gov/div898/handbook/prc/section3/prc33.htm (retrieved 2026-10-02)
- Visitors in A 1,000, Conversions in A 100, Visitors in B 1,000, Conversions in B 130, Confidence 99%, Test One-sided (B > A) gives z statistic 2.102741, p-value 0.017744, Result Not significant at 99%.Source: NIST/SEMATECH e-Handbook of Statistical Methods, §7.3.3 How can we determine whether two processes produce the same proportion of defectives? (large samples: z = (p̂₁ − p̂₂) ÷ √(p̂(1 − p̂)(1/n₁ + 1/n₂)), p̂ = (x₁ + x₂) ÷ (n₁ + n₂)), https://www.itl.nist.gov/div898/handbook/prc/section3/prc33.htm (retrieved 2026-10-02)
How it works
Group A has n_A visitors and x_A conversions; group B has n_B and x_B.
- Rates p_A = x_A ÷ n_A and p_B = x_B ÷ n_B.
- Pooled rate p̂ = (x_A + x_B) ÷ (n_A + n_B).
- Standard error SE = √(p̂ (1 − p̂) (1/n_A + 1/n_B)).
- z = (p_B − p_A) ÷ SE. A positive z means B converts better. (NIST writes p̂₁ − p̂₂ with A first, which flips the sign only.)
- p-value: two-sided p = 2 Φ(−|z|); one-sided (B better than A) p = Φ(−z), where Φ is the standard normal cumulative distribution.
- α = 1 − confidence: 0.10, 0.05 or 0.01.
- Result: "Significant at 95%: B converts better" (or "worse") when p < α; "Not significant at 95%" otherwise.
Also shown: the relative uplift (p_B − p_A) ÷ p_A × 100% (left out when p_A is 0), the difference p_B − p_A in percentage points, and the pooled rate.
Rules
- Visitors are whole numbers from 1 to 10¹²; conversions are whole numbers from 0 up to that group's visitors. More conversions than visitors has no answer.
- When no one converted in either group, or everyone did (p̂ = 0 or 1), there is no test and the page says so.
- This is the large-sample test. NIST advises Fisher's exact test for small samples.
Output format. Rates, uplift and the pooled rate are percents with 2 decimals; z and the difference have 4 decimals; the p-value has 4 significant digits.
Worked examples by hand
NIST's example: 32 of 38 against 39 of 44. p_A = 0.84211, p_B = 0.88636, p̂ = 71 ÷ 82 = 0.86585. SE = √(0.86585 × 0.13415 × (1/38 + 1/44)) = 0.075475. z = 0.044258 ÷ 0.075475 = 0.5864 (Dataplot: −0.58640 for A − B). Two-sided p = 0.5576: not significant at 95%.
200 of 10,000 against 250 of 10,000. p_A = 0.02, p_B = 0.025, p̂ = 0.0225. SE = √(0.0225 × 0.9775 × 0.0002) = 0.0020973. z = 0.005 ÷ 0.0020973 = 2.384; p = 2 Φ(−2.384) = 0.0171 < 0.05: significant at 95%, B converts better. Uplift = 0.005 ÷ 0.02 = 25%.
100 of 1,000 against 130 of 1,000, one-sided at 99%. p̂ = 0.115; SE = √(0.115 × 0.885 × 0.002) = 0.014267; z = 0.03 ÷ 0.014267 = 2.103; p = Φ(−2.103) = 0.0177, above 0.01: not significant at 99%.
Other questions people ask
How do I know if my A/B test is statistically significant?
Compute the p-value of the difference in conversion rates and compare it with 1 minus your confidence level. At 95% confidence the result is significant when p is below 0.05. With 200 of 10,000 against 250 of 10,000, z = 2.38 and p = 0.017, so B's higher rate is significant at 95% but not at 99%.
What does the p-value mean here?
It is the chance of seeing a difference at least this large if A and B truly converted at the same rate. A small p-value means the difference is unlikely to be luck alone. It is not the chance that B is better.
Should I use a one-sided or two-sided test?
Use two-sided unless you decided before the test that you only care whether B beats A. A one-sided test halves the p-value when B is ahead, so choosing it after seeing the data overstates the evidence.
Why is my test not significant even though B has more conversions?
With few visitors, chance alone moves conversion rates a lot. NIST's example of 32 of 38 against 39 of 44 is a 5% relative uplift, but z is only 0.59 and p = 0.56. More visitors make the same difference easier to detect.
What formula does this use?
The large-sample test that two proportions are equal: the pooled rate p̂ = all conversions ÷ all visitors, then z = (rate B − rate A) ÷ √(p̂(1 − p̂)(1/n_A + 1/n_B)), and the p-value from the normal distribution.
When is the z-test not accurate?
When a group has very few conversions or very few non-conversions, the normal approximation is poor. Then an exact test, such as Fisher's exact test, is the better choice.