acalculator

Did I reach statistical significance?

Type the visitors and conversions of version A and version B. The statistical significance calculator compares the two conversion rates with a two-proportion z-test and gives the z statistic, the p-value, the uplift and whether the difference is significant at 90%, 95% or 99% confidence.

Your numbers

Confidence
Test
Result
Significant at 95%: B converts better

With 200 of 10,000 against 250 of 10,000, p = 0.01713: Significant at 95%: B converts better.

p-value
0.01713
z statistic
2.384
Conversion rate A
2%
Conversion rate B
2.5%
Relative uplift
25%
Difference (points)
0.5
Pooled rate
2.25%
Significance level α
0.05

Result: Significant at 95%: B converts better. With 200 of 10,000 against 250 of 10,000, p = 0.01713: Significant at 95%: B converts better.

Where does your z fall?

How to calculate

Tests an A/B result for statistical significance: two conversion rates compared with the two-proportion z-test, giving the z statistic, the p-value and whether the difference is significant at 90%, 95% or 99%.

Example with the default inputs (Visitors in A 10,000, Conversions in A 200, Visitors in B 10,000, Conversions in B 250, Confidence 95%, Test Two-sided (A ≠ B)): With 200 of 10,000 against 250 of 10,000, p = 0.01713: Significant at 95%: B converts better.

Method: p̂ = (x_A + x_B) ÷ (n_A + n_B); z = (x_B/n_B − x_A/n_A) ÷ √(p̂(1 − p̂)(1/n_A + 1/n_B)); p = 2Φ(−|z|) two-sided or Φ(−z) one-sided; significant when p < 1 − confidence.

  • A large-sample test: the normal approximation holds when each group has several conversions and several non-conversions.
  • Visitors are independent and each is counted in one group only.
  • The one-sided test asks only whether B is better than A.

Machine-readable copies: Markdown, JSON.

Worked examples

Each example is checked against the calculator on every build.

  1. Visitors in A 38, Conversions in A 32, Visitors in B 44, Conversions in B 39, Confidence 95%, Test Two-sided (A ≠ B) gives z statistic 0.586401, p-value 0.557606, Result Not significant at 95%.Source: NIST Dataplot Reference Manual, Binomial Proportion Test (32 of 38 against 39 of 44: pooled 0.86585, statistic −0.58640, two-tailed p-value 0.55760), https://itl.nist.gov/div898/software/dataplot/refman1/auxillar/binotest.htm (retrieved 2026-10-02); NIST/SEMATECH e-Handbook of Statistical Methods, §7.3.3 How can we determine whether two processes produce the same proportion of defectives? (large samples: z = (p̂₁ − p̂₂) ÷ √(p̂(1 − p̂)(1/n₁ + 1/n₂)), p̂ = (x₁ + x₂) ÷ (n₁ + n₂)), https://www.itl.nist.gov/div898/handbook/prc/section3/prc33.htm (retrieved 2026-10-02)
  2. Visitors in A 10,000, Conversions in A 200, Visitors in B 10,000, Conversions in B 250, Confidence 95%, Test Two-sided (A ≠ B) gives z statistic 2.383995, p-value 0.017126, Relative uplift 25%, Result Significant at 95%: B converts better.Source: NIST/SEMATECH e-Handbook of Statistical Methods, §7.3.3 How can we determine whether two processes produce the same proportion of defectives? (large samples: z = (p̂₁ − p̂₂) ÷ √(p̂(1 − p̂)(1/n₁ + 1/n₂)), p̂ = (x₁ + x₂) ÷ (n₁ + n₂)), https://www.itl.nist.gov/div898/handbook/prc/section3/prc33.htm (retrieved 2026-10-02)
  3. Visitors in A 1,000, Conversions in A 100, Visitors in B 1,000, Conversions in B 130, Confidence 99%, Test One-sided (B > A) gives z statistic 2.102741, p-value 0.017744, Result Not significant at 99%.Source: NIST/SEMATECH e-Handbook of Statistical Methods, §7.3.3 How can we determine whether two processes produce the same proportion of defectives? (large samples: z = (p̂₁ − p̂₂) ÷ √(p̂(1 − p̂)(1/n₁ + 1/n₂)), p̂ = (x₁ + x₂) ÷ (n₁ + n₂)), https://www.itl.nist.gov/div898/handbook/prc/section3/prc33.htm (retrieved 2026-10-02)

How it works

Group A has n_A visitors and x_A conversions; group B has n_B and x_B.

  1. Rates p_A = x_A ÷ n_A and p_B = x_B ÷ n_B.
  2. Pooled rate p̂ = (x_A + x_B) ÷ (n_A + n_B).
  3. Standard error SE = √(p̂ (1 − p̂) (1/n_A + 1/n_B)).
  4. z = (p_B − p_A) ÷ SE. A positive z means B converts better. (NIST writes p̂₁ − p̂₂ with A first, which flips the sign only.)
  5. p-value: two-sided p = 2 Φ(−|z|); one-sided (B better than A) p = Φ(−z), where Φ is the standard normal cumulative distribution.
  6. α = 1 − confidence: 0.10, 0.05 or 0.01.
  7. Result: "Significant at 95%: B converts better" (or "worse") when p < α; "Not significant at 95%" otherwise.

Also shown: the relative uplift (p_B − p_A) ÷ p_A × 100% (left out when p_A is 0), the difference p_B − p_A in percentage points, and the pooled rate.

Rules

  • Visitors are whole numbers from 1 to 10¹²; conversions are whole numbers from 0 up to that group's visitors. More conversions than visitors has no answer.
  • When no one converted in either group, or everyone did (p̂ = 0 or 1), there is no test and the page says so.
  • This is the large-sample test. NIST advises Fisher's exact test for small samples.

Output format. Rates, uplift and the pooled rate are percents with 2 decimals; z and the difference have 4 decimals; the p-value has 4 significant digits.

Worked examples by hand

NIST's example: 32 of 38 against 39 of 44. p_A = 0.84211, p_B = 0.88636, p̂ = 71 ÷ 82 = 0.86585. SE = √(0.86585 × 0.13415 × (1/38 + 1/44)) = 0.075475. z = 0.044258 ÷ 0.075475 = 0.5864 (Dataplot: −0.58640 for A − B). Two-sided p = 0.5576: not significant at 95%.

200 of 10,000 against 250 of 10,000. p_A = 0.02, p_B = 0.025, p̂ = 0.0225. SE = √(0.0225 × 0.9775 × 0.0002) = 0.0020973. z = 0.005 ÷ 0.0020973 = 2.384; p = 2 Φ(−2.384) = 0.0171 < 0.05: significant at 95%, B converts better. Uplift = 0.005 ÷ 0.02 = 25%.

100 of 1,000 against 130 of 1,000, one-sided at 99%. p̂ = 0.115; SE = √(0.115 × 0.885 × 0.002) = 0.014267; z = 0.03 ÷ 0.014267 = 2.103; p = Φ(−2.103) = 0.0177, above 0.01: not significant at 99%.

Other questions people ask

How do I know if my A/B test is statistically significant?

Compute the p-value of the difference in conversion rates and compare it with 1 minus your confidence level. At 95% confidence the result is significant when p is below 0.05. With 200 of 10,000 against 250 of 10,000, z = 2.38 and p = 0.017, so B's higher rate is significant at 95% but not at 99%.

What does the p-value mean here?

It is the chance of seeing a difference at least this large if A and B truly converted at the same rate. A small p-value means the difference is unlikely to be luck alone. It is not the chance that B is better.

Should I use a one-sided or two-sided test?

Use two-sided unless you decided before the test that you only care whether B beats A. A one-sided test halves the p-value when B is ahead, so choosing it after seeing the data overstates the evidence.

Why is my test not significant even though B has more conversions?

With few visitors, chance alone moves conversion rates a lot. NIST's example of 32 of 38 against 39 of 44 is a 5% relative uplift, but z is only 0.59 and p = 0.56. More visitors make the same difference easier to detect.

What formula does this use?

The large-sample test that two proportions are equal: the pooled rate p̂ = all conversions ÷ all visitors, then z = (rate B − rate A) ÷ √(p̂(1 − p̂)(1/n_A + 1/n_B)), and the p-value from the normal distribution.

When is the z-test not accurate?

When a group has very few conversions or very few non-conversions, the normal approximation is poor. Then an exact test, such as Fisher's exact test, is the better choice.