acalculator

What is my test statistic?

Pick a test and type your summary numbers. The test statistic calculator shows the z or t statistic, its standard error, the degrees of freedom for t, and the p-value for a two-tailed, left or right test.

Your numbers

Alternative
Test statistic
1.7817

The test statistic is t = 1.7817, with a p-value of 0.1085.

Statistic
t
Standard error
2.07667
Degrees of freedom
9
p-value
0.1085
Estimate
53.7

Test statistic: 1.7817. The test statistic is t = 1.7817, with a p-value of 0.1085.

Where does your statistic fall?

How to calculate

Computes the z or t test statistic, its standard error, degrees of freedom and p-value for one mean, two means, one proportion or two proportions, from summary numbers.

Example with the default inputs (Test One mean, σ unknown (t), Alternative Two-tailed (≠), Sample mean x̄₁ 53.7, Hypothesized mean μ₀ 50, Standard deviation 6.567, Sample size n₁ 10): The test statistic is t = 1.7817, with a p-value of 0.1085.

Method: z = (x̄ − μ₀) ÷ (σ ÷ √n); t = (x̄ − μ₀) ÷ (s ÷ √n), n − 1 df; two means: t = (x̄₁ − x̄₂) ÷ √(s₁²/n₁ + s₂²/n₂) (Welch df) or ÷ (sₚ√(1/n₁ + 1/n₂)) (pooled, n₁ + n₂ − 2 df); one proportion: z = (p̂ − p₀) ÷ √(p₀(1 − p₀)/n); two proportions: z = (p̂₁ − p̂₂) ÷ √(p̂(1 − p̂)(1/n₁ + 1/n₂)).

  • Random, independent samples. t tests assume roughly normal data or large samples; z tests for proportions use the normal approximation, which needs enough successes and failures (often 10 or more of each).
  • The two-means test compares x̄₁ − x̄₂ with 0.
  • p-values: two-tailed = 2 × the upper tail beyond |statistic| (at most 1); left = the lower tail below it; right = the upper tail above it.

Machine-readable copies: Markdown, JSON.

Worked examples

Each example is checked against the calculator on every build.

  1. Test One mean, σ unknown (t), Alternative Two-tailed (≠), Sample mean x̄₁ 53.7, Hypothesized mean μ₀ 50, Standard deviation 6.567, Sample size n₁ 10 gives Test statistic 1.781701, Degrees of freedom 9.Source: NIST/SEMATECH e-Handbook of Statistical Methods, §7.2.2 (t = (Ȳ − μ₀) ÷ (s ÷ √N), N − 1 df; 10 wafers, mean 53.7, s 6.567, μ₀ 50: t = 1.782). https://www.itl.nist.gov/div898/handbook/prc/section2/prc22.htm
  2. Test Two means (t), Alternative Two-tailed (≠), Variances Equal (pooled), Sample mean x̄₁ 20.14458, Standard deviation 6.4147, Sample size n₁ 249, Sample mean x̄₂ 30.48101, Standard deviation s₂ 6.10771, Sample size n₂ 79 gives Test statistic -12.620585, Degrees of freedom 326.Source: NIST/SEMATECH e-Handbook of Statistical Methods, §1.3.5.3 Two-Sample t-Test (AUTO83B.DAT: N 249 and 79, means 20.14458 and 30.48101, SD 6.41470 and 6.10771; pooled T = −12.62059, 326 df). https://www.itl.nist.gov/div898/handbook/eda/section3/eda353.htm
  3. Test One proportion (z), Alternative Right (>), Successes x₁ 26, Sample size n₁ 200, Hypothesized proportion p₀ 0.1 gives Test statistic 1.414214, Estimate 0.13, p-value 0.07865.Source: NIST/SEMATECH e-Handbook of Statistical Methods, §7.2.4 (z = (p̂ − p₀) ÷ √(p₀(1 − p₀) ÷ N); 26 of 200, p₀ 0.10: z = 1.414). https://www.itl.nist.gov/div898/handbook/prc/section2/prc24.htm
  4. Test One mean, σ known (z), Alternative Two-tailed (≠), Sample mean x̄₁ 103, Hypothesized mean μ₀ 100, Standard deviation 15, Sample size n₁ 25 gives Test statistic 1, Standard error 3, p-value 0.317311.Source: NIST/SEMATECH e-Handbook of Statistical Methods, §7.2.2 (z = (Ȳ − μ₀) ÷ (σ ÷ √N)). https://www.itl.nist.gov/div898/handbook/prc/section2/prc22.htm
  5. Test Two proportions (z), Alternative Two-tailed (≠), Successes x₁ 45, Sample size n₁ 100, Successes x₂ 30, Sample size n₂ 100 gives Test statistic 2.19089, Estimate 0.15.Source: NIST/SEMATECH e-Handbook of Statistical Methods, §7.3.3 (z = (p̂₁ − p̂₂) ÷ √(p̂(1 − p̂)(1/n₁ + 1/n₂)), p̂ = (x₁ + x₂) ÷ (n₁ + n₂)). https://www.itl.nist.gov/div898/handbook/prc/section3/prc33.htm

How it works

TestEstimateStandard error (SE)StatisticDistribution
One mean, σ knownx̄σ ÷ √nz = (x̄ − μ₀) ÷ SEstandard normal
One mean, σ unknownx̄s ÷ √nt = (x̄ − μ₀) ÷ SEt, n − 1 df
Two means, Welchx̄₁ − x̄₂√(s₁²/n₁ + s₂²/n₂)t = (x̄₁ − x̄₂) ÷ SEt, Welch df
Two means, pooledx̄₁ − x̄₂sₚ √(1/n₁ + 1/n₂)t = (x̄₁ − x̄₂) ÷ SEt, n₁ + n₂ − 2 df
One proportionp̂ = x ÷ n√(p₀(1 − p₀) ÷ n)z = (p̂ − p₀) ÷ SEstandard normal
Two proportionsp̂₁ − p̂₂√(p̂(1 − p̂)(1/n₁ + 1/n₂))z = (p̂₁ − p̂₂) ÷ SEstandard normal
  • Pooled variance: sₚ² = ((n₁ − 1)s₁² + (n₂ − 1)s₂²) ÷ (n₁ + n₂ − 2)
  • Welch–Satterthwaite: df = (s₁²/n₁ + s₂²/n₂)² ÷ ((s₁²/n₁)²/(n₁ − 1) + (s₂²/n₂)²/(n₂ − 1)), not rounded
  • Pooled proportion: p̂ = (x₁ + x₂) ÷ (n₁ + n₂)
  • p-value: right-tailed = P(X > statistic); left-tailed = P(X < statistic); two-tailed = 2 × P(X > |statistic|), at most 1. X is standard normal for z and Student’s t with the degrees of freedom for t (tail areas precise at any df).

The chart shades the p-value area under the normal or t curve and marks the statistic.

Arithmetic. 64-bit floats throughout.

Output format. The statistic and the degrees of freedom show 4 decimals, the standard error 6 significant figures, the p-value 4 significant figures and the estimate 8, each rounded half up, with trailing zeros left off.

When there is no answer.

  • A t test with fewer than 2 observations in a sample.
  • More successes than the sample size.
  • Two proportions with no successes, or only successes, in both samples (the pooled SE is 0).
  • Numbers that give a standard error of 0 or a statistic beyond the largest number the page can hold.

Assumptions

  • Random, independent samples. The t tests assume roughly normal data or large samples; the proportion tests use the normal approximation, which needs enough successes and failures (often 10 or more of each).
  • Standard deviations are more than 0; sample sizes are whole numbers from 1 to 10⁹; p₀ is strictly between 0 and 1.

Worked examples by hand

NIST wafers, one mean, t. SE = 6.567 ÷ √10 = 2.07666; t = (53.7 − 50) ÷ 2.07666 = 1.7817, 9 df.

NIST car mileage, two means, pooled. sₚ² = (248 × 6.41470² + 78 × 6.10771²) ÷ 326 = 40.229; SE = √(40.229 × (1/249 + 1/79)) = 0.81901; t = (20.14458 − 30.48101) ÷ 0.81901 = −12.6206, 326 df.

NIST wafer defects, one proportion, right-tailed. p̂ = 26 ÷ 200 = 0.13; SE = √(0.1 × 0.9 ÷ 200) = 0.021213; z = 0.03 ÷ 0.021213 = 1.4142; p = P(Z > 1.4142) = 0.07865.

One mean, σ known. SE = 15 ÷ √25 = 3; z = (103 − 100) ÷ 3 = 1; two-tailed p = 2 × (1 − Φ(1)) = 0.3173.

Two proportions. 45 of 100 and 30 of 100: estimate 0.15; p̂ = 75 ÷ 200 = 0.375; SE = √(0.375 × 0.625 × 0.02) = 0.068465; z = 2.1909.

Other questions people ask

What is a test statistic?

It measures how far your sample estimate is from the value in the null hypothesis, in standard errors: (estimate − null value) ÷ standard error. A large statistic in either direction is evidence against the null hypothesis.

How do I calculate a t statistic for one mean?

t = (x̄ − μ₀) ÷ (s ÷ √n), with n − 1 degrees of freedom. NIST’s example has 10 wafers with mean 53.7 and s = 6.567 against μ₀ = 50: t = 3.7 ÷ 2.0767 = 1.782, with 9 degrees of freedom.

When do I use z instead of t?

Use z for a mean when the population standard deviation σ is known, and for proportions (with large samples). Use t when you estimate the standard deviation from the sample.

How do I calculate a z statistic for a proportion?

z = (p̂ − p₀) ÷ √(p₀(1 − p₀) ÷ n). NIST’s example: 26 defective wafers out of 200 is p̂ = 0.13; against p₀ = 0.10, z = 0.03 ÷ 0.02121 = 1.414.

Should I use Welch’s or the pooled two-sample t test?

Welch’s test does not assume the two groups have the same variance and is the safer default. The pooled test assumes equal variances and uses n₁ + n₂ − 2 degrees of freedom. NIST’s car mileage example uses the pooled test: T = −12.62 with 326 degrees of freedom.

How do I get the p-value from the test statistic?

For a right-tailed test, the p-value is the area above the statistic; for a left-tailed test, the area below; for a two-tailed test, twice the area beyond |statistic|. The page uses the standard normal for z and Student’s t with the degrees of freedom for t.