# Did I reach statistical significance?

Tests an A/B result for statistical significance: two conversion rates compared with the two-proportion z-test, giving the z statistic, the p-value and whether the difference is significant at 90%, 95% or 99%.

- Page: https://www.acalculator.org/statistics/statistical-significance-calculator
- JSON spec: https://www.acalculator.org/statistics/statistical-significance-calculator.json
- Version: ab69d0308b34

## Default answer

Example with the default inputs (Visitors in A 10,000, Conversions in A 200, Visitors in B 10,000, Conversions in B 250, Confidence 95%, Test Two-sided (A ≠ B)): With 200 of 10,000 against 250 of 10,000, p = 0.01713: Significant at 95%: B converts better.

## Inputs

| Key | Label | Description |
| --- | --- | --- |
| na | Visitors in A | How many people saw version A (the control). |
| xa | Conversions in A | How many of them converted, clicked or bought. |
| nb | Visitors in B | How many people saw version B (the variant). |
| xb | Conversions in B | How many of them converted. |
| conf | Confidence | How sure you want to be: significant when the p-value is below 1 minus this. |
| tail | Test | Two-sided asks whether A and B differ; one-sided asks whether B is better than A. |

## Outputs

| Key | Label | Description |
| --- | --- | --- |
| verdict | Result | Whether the difference is statistically significant. |
| p | p-value | The chance of a difference this large or larger if A and B truly convert at the same rate. |
| z | z statistic | (rate B − rate A) ÷ the pooled standard error. |
| rateA | Conversion rate A | Conversions in A ÷ visitors in A. |
| rateB | Conversion rate B | Conversions in B ÷ visitors in B. |
| uplift | Relative uplift | (rate B − rate A) ÷ rate A: how much better or worse B does. |
| diff | Difference (points) | Rate B − rate A, in percentage points. |
| pooled | Pooled rate | All conversions ÷ all visitors. |
| alpha | Significance level α | 1 − confidence: the p-value must be below it. |

## Method

p̂ = (x_A + x_B) ÷ (n_A + n_B); z = (x_B/n_B − x_A/n_A) ÷ √(p̂(1 − p̂)(1/n_A + 1/n_B)); p = 2Φ(−|z|) two-sided or Φ(−z) one-sided; significant when p < 1 − confidence.

## Assumptions

- A large-sample test: the normal approximation holds when each group has several conversions and several non-conversions.
- Visitors are independent and each is counted in one group only.
- The one-sided test asks only whether B is better than A.

## Worked examples

1. na = 38, xa = 32, nb = 44, xb = 39, conf = 95, tail = two gives z = 0.586401, p = 0.557606, verdict = Not significant at 95%. Source: NIST Dataplot Reference Manual, Binomial Proportion Test (32 of 38 against 39 of 44: pooled 0.86585, statistic −0.58640, two-tailed p-value 0.55760), https://itl.nist.gov/div898/software/dataplot/refman1/auxillar/binotest.htm (retrieved 2026-10-02); NIST/SEMATECH e-Handbook of Statistical Methods, §7.3.3 How can we determine whether two processes produce the same proportion of defectives? (large samples: z = (p̂₁ − p̂₂) ÷ √(p̂(1 − p̂)(1/n₁ + 1/n₂)), p̂ = (x₁ + x₂) ÷ (n₁ + n₂)), https://www.itl.nist.gov/div898/handbook/prc/section3/prc33.htm (retrieved 2026-10-02).
2. na = 10,000, xa = 200, nb = 10,000, xb = 250, conf = 95, tail = two gives z = 2.383995, p = 0.017126, uplift = 25%, verdict = Significant at 95%: B converts better. Source: NIST/SEMATECH e-Handbook of Statistical Methods, §7.3.3 How can we determine whether two processes produce the same proportion of defectives? (large samples: z = (p̂₁ − p̂₂) ÷ √(p̂(1 − p̂)(1/n₁ + 1/n₂)), p̂ = (x₁ + x₂) ÷ (n₁ + n₂)), https://www.itl.nist.gov/div898/handbook/prc/section3/prc33.htm (retrieved 2026-10-02).
3. na = 1,000, xa = 100, nb = 1,000, xb = 130, conf = 99, tail = one gives z = 2.102741, p = 0.017744, verdict = Not significant at 99%. Source: NIST/SEMATECH e-Handbook of Statistical Methods, §7.3.3 How can we determine whether two processes produce the same proportion of defectives? (large samples: z = (p̂₁ − p̂₂) ÷ √(p̂(1 − p̂)(1/n₁ + 1/n₂)), p̂ = (x₁ + x₂) ÷ (n₁ + n₂)), https://www.itl.nist.gov/div898/handbook/prc/section3/prc33.htm (retrieved 2026-10-02).

## FAQ

### How do I know if my A/B test is statistically significant?

Compute the p-value of the difference in conversion rates and compare it with 1 minus your confidence level. At 95% confidence the result is significant when p is below 0.05. With 200 of 10,000 against 250 of 10,000, z = 2.38 and p = 0.017, so B's higher rate is significant at 95% but not at 99%.

### What does the p-value mean here?

It is the chance of seeing a difference at least this large if A and B truly converted at the same rate. A small p-value means the difference is unlikely to be luck alone. It is not the chance that B is better.

### Should I use a one-sided or two-sided test?

Use two-sided unless you decided before the test that you only care whether B beats A. A one-sided test halves the p-value when B is ahead, so choosing it after seeing the data overstates the evidence.

### Why is my test not significant even though B has more conversions?

With few visitors, chance alone moves conversion rates a lot. NIST's example of 32 of 38 against 39 of 44 is a 5% relative uplift, but z is only 0.59 and p = 0.56. More visitors make the same difference easier to detect.

### What formula does this use?

The large-sample test that two proportions are equal: the pooled rate p̂ = all conversions ÷ all visitors, then z = (rate B − rate A) ÷ √(p̂(1 − p̂)(1/n_A + 1/n_B)), and the p-value from the normal distribution.

### When is the z-test not accurate?

When a group has very few conversions or very few non-conversions, the normal approximation is poor. Then an exact test, such as Fisher's exact test, is the better choice.

## Sources

- NIST/SEMATECH e-Handbook of Statistical Methods, §7.3.3 How can we determine whether two processes produce the same proportion of defectives? (Case 1, large samples: z = (p̂₁ − p̂₂) ÷ √(p̂(1 − p̂)(1/n₁ + 1/n₂)) with p̂ = (x₁ + x₂) ÷ (n₁ + n₂); Case 2, small samples: Fisher's exact test). https://www.itl.nist.gov/div898/handbook/prc/section3/prc33.htm (retrieved 2026-10-02)
- NIST Dataplot Reference Manual, Binomial Proportion Test (critical regions for two-tailed, lower- and upper-tailed tests; p-value 2(1 − Φ(|Z|)) two-tailed; example: 32 of 38 against 39 of 44, pooled 0.86585, Z = −0.58640, two-tailed p-value 0.55760). https://itl.nist.gov/div898/software/dataplot/refman1/auxillar/binotest.htm (retrieved 2026-10-02)
