{
  "id": "statistical-significance",
  "version": "ab69d0308b34",
  "status": "published",
  "name": "Statistical Significance Calculator",
  "question": "Did I reach statistical significance?",
  "summary": "Tests an A/B result for statistical significance: two conversion rates compared with the two-proportion z-test, giving the z statistic, the p-value and whether the difference is significant at 90%, 95% or 99%.",
  "category": "statistics",
  "subcategory": "tests",
  "url": "https://www.acalculator.org/statistics/statistical-significance-calculator",
  "markdown": "https://www.acalculator.org/statistics/statistical-significance-calculator.md",
  "kind": "function",
  "method": "p̂ = (x_A + x_B) ÷ (n_A + n_B); z = (x_B/n_B − x_A/n_A) ÷ √(p̂(1 − p̂)(1/n_A + 1/n_B)); p = 2Φ(−|z|) two-sided or Φ(−z) one-sided; significant when p < 1 − confidence.",
  "assumptions": [
    "A large-sample test: the normal approximation holds when each group has several conversions and several non-conversions.",
    "Visitors are independent and each is counted in one group only.",
    "The one-sided test asks only whether B is better than A."
  ],
  "inputs": {
    "$schema": "https://json-schema.org/draft/2020-12/schema",
    "type": "object",
    "properties": {
      "na": {
        "title": "Visitors in A",
        "description": "How many people saw version A (the control).",
        "type": "integer",
        "minimum": 1,
        "maximum": 1000000000000
      },
      "xa": {
        "title": "Conversions in A",
        "description": "How many of them converted, clicked or bought.",
        "type": "integer",
        "minimum": 0,
        "maximum": 1000000000000
      },
      "nb": {
        "title": "Visitors in B",
        "description": "How many people saw version B (the variant).",
        "type": "integer",
        "minimum": 1,
        "maximum": 1000000000000
      },
      "xb": {
        "title": "Conversions in B",
        "description": "How many of them converted.",
        "type": "integer",
        "minimum": 0,
        "maximum": 1000000000000
      },
      "conf": {
        "title": "Confidence",
        "description": "How sure you want to be: significant when the p-value is below 1 minus this.",
        "type": "string",
        "enum": [
          "90",
          "95",
          "99"
        ]
      },
      "tail": {
        "title": "Test",
        "description": "Two-sided asks whether A and B differ; one-sided asks whether B is better than A.",
        "type": "string",
        "enum": [
          "two",
          "one"
        ]
      }
    }
  },
  "outputs": {
    "verdict": {
      "label": "Result",
      "description": "Whether the difference is statistically significant.",
      "format": "text"
    },
    "p": {
      "label": "p-value",
      "description": "The chance of a difference this large or larger if A and B truly convert at the same rate.",
      "format": "number"
    },
    "z": {
      "label": "z statistic",
      "description": "(rate B − rate A) ÷ the pooled standard error.",
      "format": "number"
    },
    "rateA": {
      "label": "Conversion rate A",
      "description": "Conversions in A ÷ visitors in A.",
      "format": "percent"
    },
    "rateB": {
      "label": "Conversion rate B",
      "description": "Conversions in B ÷ visitors in B.",
      "format": "percent"
    },
    "uplift": {
      "label": "Relative uplift",
      "description": "(rate B − rate A) ÷ rate A: how much better or worse B does.",
      "format": "percent"
    },
    "diff": {
      "label": "Difference (points)",
      "description": "Rate B − rate A, in percentage points.",
      "format": "number"
    },
    "pooled": {
      "label": "Pooled rate",
      "description": "All conversions ÷ all visitors.",
      "format": "percent"
    },
    "alpha": {
      "label": "Significance level α",
      "description": "1 − confidence: the p-value must be below it.",
      "format": "number"
    }
  },
  "defaultAnswer": {
    "inputs": {
      "na": 10000,
      "xa": 200,
      "nb": 10000,
      "xb": 250,
      "conf": "95",
      "tail": "two"
    },
    "outputs": {
      "verdict": "Significant at 95%: B converts better",
      "p": 0.017125828872195006,
      "z": 2.3839951328405053,
      "rateA": 2,
      "rateB": 2.5,
      "uplift": 25.000000000000007,
      "diff": 0.5000000000000001,
      "pooled": 2.25,
      "alpha": 0.05
    },
    "text": "With 200 of 10,000 against 250 of 10,000, p = 0.01713: Significant at 95%: B converts better."
  },
  "examples": [
    {
      "given": {
        "na": 38,
        "xa": 32,
        "nb": 44,
        "xb": 39,
        "conf": "95",
        "tail": "two"
      },
      "expect": {
        "z": 0.5864013918242953,
        "p": 0.5576058092182463,
        "verdict": "Not significant at 95%"
      },
      "source": "NIST Dataplot Reference Manual, Binomial Proportion Test (32 of 38 against 39 of 44: pooled 0.86585, statistic −0.58640, two-tailed p-value 0.55760), https://itl.nist.gov/div898/software/dataplot/refman1/auxillar/binotest.htm (retrieved 2026-10-02); NIST/SEMATECH e-Handbook of Statistical Methods, §7.3.3 How can we determine whether two processes produce the same proportion of defectives? (large samples: z = (p̂₁ − p̂₂) ÷ √(p̂(1 − p̂)(1/n₁ + 1/n₂)), p̂ = (x₁ + x₂) ÷ (n₁ + n₂)), https://www.itl.nist.gov/div898/handbook/prc/section3/prc33.htm (retrieved 2026-10-02); z here is B − A, so its sign is the opposite of Dataplot’s −0.58640",
      "tolerance": 1e-9
    },
    {
      "given": {
        "na": 10000,
        "xa": 200,
        "nb": 10000,
        "xb": 250,
        "conf": "95",
        "tail": "two"
      },
      "expect": {
        "z": 2.3839951328405053,
        "p": 0.01712582887219502,
        "uplift": 25.000000000000007,
        "verdict": "Significant at 95%: B converts better"
      },
      "source": "NIST/SEMATECH e-Handbook of Statistical Methods, §7.3.3 How can we determine whether two processes produce the same proportion of defectives? (large samples: z = (p̂₁ − p̂₂) ÷ √(p̂(1 − p̂)(1/n₁ + 1/n₂)), p̂ = (x₁ + x₂) ÷ (n₁ + n₂)), https://www.itl.nist.gov/div898/handbook/prc/section3/prc33.htm (retrieved 2026-10-02); hand calculation in content.mdx: p̂ = 450 ÷ 20,000 = 0.0225; SE = √(0.0225 × 0.9775 × 0.0002) = 0.0020973; z = 0.005 ÷ 0.0020973 = 2.384; p = 0.0171; Python 3 math.erfc",
      "tolerance": 1e-9
    },
    {
      "given": {
        "na": 1000,
        "xa": 100,
        "nb": 1000,
        "xb": 130,
        "conf": "99",
        "tail": "one"
      },
      "expect": {
        "z": 2.102740605622114,
        "p": 0.017744225233237366,
        "verdict": "Not significant at 99%"
      },
      "source": "NIST/SEMATECH e-Handbook of Statistical Methods, §7.3.3 How can we determine whether two processes produce the same proportion of defectives? (large samples: z = (p̂₁ − p̂₂) ÷ √(p̂(1 − p̂)(1/n₁ + 1/n₂)), p̂ = (x₁ + x₂) ÷ (n₁ + n₂)), https://www.itl.nist.gov/div898/handbook/prc/section3/prc33.htm (retrieved 2026-10-02); hand calculation in content.mdx: p̂ = 0.115; SE = 0.014267; z = 0.03 ÷ 0.014267 = 2.103; one-sided p = 0.0177 > 0.01",
      "tolerance": 1e-9
    }
  ],
  "sources": [
    "NIST/SEMATECH e-Handbook of Statistical Methods, §7.3.3 How can we determine whether two processes produce the same proportion of defectives? (Case 1, large samples: z = (p̂₁ − p̂₂) ÷ √(p̂(1 − p̂)(1/n₁ + 1/n₂)) with p̂ = (x₁ + x₂) ÷ (n₁ + n₂); Case 2, small samples: Fisher's exact test). https://www.itl.nist.gov/div898/handbook/prc/section3/prc33.htm (retrieved 2026-10-02)",
    "NIST Dataplot Reference Manual, Binomial Proportion Test (critical regions for two-tailed, lower- and upper-tailed tests; p-value 2(1 − Φ(|Z|)) two-tailed; example: 32 of 38 against 39 of 44, pooled 0.86585, Z = −0.58640, two-tailed p-value 0.55760). https://itl.nist.gov/div898/software/dataplot/refman1/auxillar/binotest.htm (retrieved 2026-10-02)"
  ],
  "related": [
    "p-value",
    "z-score",
    "sample-size",
    "confidence-interval"
  ],
  "changelog": []
}
