# Hypothesis testing: do I reject H₀?

Runs a z or t hypothesis test for one mean, two means, one proportion or two proportions from summary numbers, and compares the p-value with your significance level to reject or keep H₀.

- Page: https://www.acalculator.org/statistics/hypothesis-test-calculator
- JSON spec: https://www.acalculator.org/statistics/hypothesis-test-calculator.json
- Version: 50972ea1c8c0

## Default answer

Example with the default inputs (Test One mean, σ unknown (t), Alternative hypothesis H₁ Two-tailed (≠), Significance level (α) 0.05, Sample mean x̄₁ 53.7, Hypothesized mean μ₀ 50, Standard deviation 6.567, Sample size n₁ 10): Do not reject H₀: the p-value is not less than α = 0.05. The test statistic is t = 1.7817, with a p-value of 0.1085.

## Inputs

| Key | Label | Description |
| --- | --- | --- |
| test | Test | Which hypothesis test to run. |
| tail | Alternative hypothesis H₁ | Two-tailed (≠), left-tailed (<) or right-tailed (>). |
| alpha | Significance level (α) | The chance of rejecting a true H₀ that you accept, often 0.05. |
| m1 | Sample mean x̄₁ | The mean of the (first) sample. |
| mu | Hypothesized mean μ₀ | The population mean under H₀. |
| sd1 | Standard deviation | σ for a z test; the sample standard deviation s for a t test. |
| n1 | Sample size n₁ | How many observations are in the (first) sample. |
| m2 | Sample mean x̄₂ | The mean of the second sample. |
| sd2 | Standard deviation s₂ | The standard deviation of the second sample. |
| n2 | Sample size n₂ | How many observations are in the second sample. |
| var | Variances | Welch’s test for unequal variances, or the pooled test for equal variances. |
| x1 | Successes x₁ | How many of the (first) sample have the trait. |
| p0 | Hypothesized proportion p₀ | The population proportion under H₀, between 0 and 1. |
| x2 | Successes x₂ | How many of the second sample have the trait. |

## Outputs

| Key | Label | Description |
| --- | --- | --- |
| decision | Decision | Reject H₀ when the p-value is less than α; otherwise do not reject H₀. |
| p | p-value | The chance of a statistic at least this extreme in the direction of H₁, if H₀ holds. |
| statistic | Test statistic | The estimate minus the null value, divided by its standard error. |
| critical | Critical value | The statistic beyond which H₀ is rejected at α, to 4 decimals. |
| kind | Statistic | z (standard normal) or t (Student’s t). |
| df | Degrees of freedom | For t: n − 1, n₁ + n₂ − 2 (pooled) or the Welch–Satterthwaite value. |
| se | Standard error | The standard error of the estimate under H₀. |
| estimate | Estimate | x̄, x̄₁ − x̄₂, p̂ or p̂₁ − p̂₂. |

## Method

Compute z or t as on the test statistic page; p-value by the tail of H₁; critical value z or t at α (two-tailed: α ÷ 2 in each tail); reject H₀ when p < α.

## Assumptions

- Random, independent samples. t tests assume roughly normal data or large samples; the proportion tests use the normal approximation.
- The two-means and two-proportions tests compare the difference with 0.
- A p-value equal to α does not reject H₀.

## Worked examples

1. test = t, tail = two, alpha = 0.05, m1 = 53.7, mu = 50, sd1 = 6.567, n1 = 10 gives decision = Do not reject H₀: the p-value is not less than α = 0.05, statistic = 1.781701, df = 9, critical = ±2.2622. Source: NIST/SEMATECH e-Handbook of Statistical Methods, §7.2.2 Are the data consistent with the assumed process mean? (t = (Ȳ − μ₀) ÷ (s ÷ √N), N − 1 df; 10 wafers, mean 53.7, s 6.567, μ₀ 50: t = 1.782, below the critical 2.262 at α = 0.05, so H₀ is not rejected), https://www.itl.nist.gov/div898/handbook/prc/section2/prc22.htm (retrieved 2026-10-05); OpenStax, Introductory Statistics 2e, §9.4 Rare Events, the Sample, and the Decision and Conclusion (reject H₀ when the p-value is less than α; otherwise do not reject H₀), https://openstax.org/books/introductory-statistics-2e/pages/9-4-rare-events-the-sample-decision-and-conclusion (retrieved 2026-10-05).
2. test = prop, tail = right, alpha = 0.05, x1 = 26, n1 = 200, p0 = 0.1 gives decision = Do not reject H₀: the p-value is not less than α = 0.05, statistic = 1.414214, p = 0.07865, critical = 1.6449, estimate = 0.13. Source: NIST/SEMATECH e-Handbook of Statistical Methods, §7.2.4 Does the proportion of defectives meet requirements? (z = (p̂ − p₀) ÷ √(p₀(1 − p₀) ÷ N); 26 of 200, p₀ 0.10: z = 1.414, below the critical 1.645 at α = 0.05), https://www.itl.nist.gov/div898/handbook/prc/section2/prc24.htm (retrieved 2026-10-05); OpenStax, Introductory Statistics 2e, §9.4 Rare Events, the Sample, and the Decision and Conclusion (reject H₀ when the p-value is less than α; otherwise do not reject H₀), https://openstax.org/books/introductory-statistics-2e/pages/9-4-rare-events-the-sample-decision-and-conclusion (retrieved 2026-10-05).
3. test = two, tail = two, alpha = 0.05, var = pooled, m1 = 20.14458, sd1 = 6.4147, n1 = 249, m2 = 30.48101, sd2 = 6.10771, n2 = 79 gives decision = Reject H₀: the p-value is less than α = 0.05, statistic = -12.620585, df = 326, critical = ±1.9673. Source: NIST/SEMATECH e-Handbook of Statistical Methods, §1.3.5.3 Two-Sample t-Test for Equal Means (pooled T = −12.62059 with 326 df; critical 1.9673 at α = 0.05, so H₀ is rejected), https://www.itl.nist.gov/div898/handbook/eda/section3/eda353.htm (retrieved 2026-10-05); OpenStax, Introductory Statistics 2e, §9.4 Rare Events, the Sample, and the Decision and Conclusion (reject H₀ when the p-value is less than α; otherwise do not reject H₀), https://openstax.org/books/introductory-statistics-2e/pages/9-4-rare-events-the-sample-decision-and-conclusion (retrieved 2026-10-05).
4. test = z, tail = left, alpha = 0.05, m1 = 97, mu = 100, sd1 = 15, n1 = 100 gives decision = Reject H₀: the p-value is less than α = 0.05, statistic = -2, p = 0.02275, critical = −1.6449. Source: OpenStax, Introductory Statistics 2e, §9.4 Rare Events, the Sample, and the Decision and Conclusion (reject H₀ when the p-value is less than α; otherwise do not reject H₀), https://openstax.org/books/introductory-statistics-2e/pages/9-4-rare-events-the-sample-decision-and-conclusion (retrieved 2026-10-05).

## FAQ

### How do I do a hypothesis test?

State H₀ and the alternative H₁, pick a significance level α (often 0.05), compute the test statistic from your sample, find its p-value, and compare. If the p-value is less than α, reject H₀; otherwise do not reject it.

### When do I reject the null hypothesis?

Reject H₀ when the p-value is less than α. Equally, reject it when the test statistic is beyond the critical value in the direction of H₁. Both rules always give the same decision.

### Does "do not reject H₀" mean H₀ is true?

No. It means the sample does not give strong enough evidence against H₀ at the chosen α. A larger sample might.

### Should I use a z test or a t test?

Use a z test for a mean when the population standard deviation σ is known, and for proportions. Use a t test for a mean when you only have the sample standard deviation s. For large samples the two give almost the same answer.

### What is a one-tailed versus a two-tailed test?

A two-tailed test (H₁: ≠) looks for a difference in either direction and splits α between the two tails. A left-tailed (<) or right-tailed (>) test looks in one direction only and puts all of α in that tail, so its critical value is closer to 0.

### What significance level should I use?

0.05 is the most common. Use 0.01 when a false rejection would be costly, or 0.10 for an early look. Choose α before you see the data.

## Sources

- NIST/SEMATECH e-Handbook of Statistical Methods, §7.2.2 Are the data consistent with the assumed process mean?: t = (Ȳ − μ₀) ÷ (s ÷ √N) with N − 1 df; the wafer example gives t = 1.782 against a critical 2.262. https://www.itl.nist.gov/div898/handbook/prc/section2/prc22.htm (retrieved 2026-10-05)
- NIST/SEMATECH e-Handbook of Statistical Methods, §7.2.4 Does the proportion of defectives meet requirements?: z = (p̂ − p₀) ÷ √(p₀(1 − p₀) ÷ N); 26 of 200 against p₀ = 0.10 gives z = 1.414. https://www.itl.nist.gov/div898/handbook/prc/section2/prc24.htm (retrieved 2026-10-05)
- NIST/SEMATECH e-Handbook of Statistical Methods, §1.3.5.3 Two-Sample t-Test for Equal Means: pooled and Welch statistics and degrees of freedom; the car mileage example gives T = −12.62 with 326 df. https://www.itl.nist.gov/div898/handbook/eda/section3/eda353.htm (retrieved 2026-10-05)
- NIST/SEMATECH e-Handbook of Statistical Methods, §7.3.3 How can we determine whether two processes produce the same proportion of defectives?: z for two proportions with the pooled p̂. https://www.itl.nist.gov/div898/handbook/prc/section3/prc33.htm (retrieved 2026-10-05)
- OpenStax, Introductory Statistics 2e, §9.4 Rare Events, the Sample, and the Decision and Conclusion: reject H₀ when the p-value is less than α; otherwise do not reject H₀. https://openstax.org/books/introductory-statistics-2e/pages/9-4-rare-events-the-sample-decision-and-conclusion (retrieved 2026-10-05)
