What is my p-value?
Find the p-value of your z, t, chi-square, or F test statistic, and what it says about the null hypothesis.
- P-value
- 0.049996
The p-value is 0.049996 for a Z statistic of 1.96 (Two-tailed).
Moderate evidence against the null hypothesis
- Very strong evidence against the null hypothesis
- Strong evidence against the null hypothesis
- Moderate evidence against the null hypothesis
- Weak evidence against the null hypothesis
- Insufficient evidence to reject the null hypothesis
- Left tail, P(X ≤ statistic)
- 0.975002
- Right tail, P(X ≥ statistic)
- 0.024998
P-value: 0.049996. The p-value is 0.049996 for a Z statistic of 1.96 (Two-tailed).
Where does your statistic fall?
How to calculate
Computes the p-value of a z, t, chi-square, or F test statistic for a left-, right-, or two-tailed test.
Example with the default inputs (Distribution Z, Test type Two-tailed, Test statistic 1.96): The p-value is 0.049996 for a Z statistic of 1.96 (Two-tailed).
Method: Left-tailed: P(X ≤ s); right-tailed: P(X ≥ s); two-tailed: 2 × the smaller of the two, at most 1, for the chosen distribution under the null hypothesis.
- For chi-square and F, which are not symmetric, the two-tailed p-value doubles the smaller tail.
- Degrees of freedom may be fractional (for example Welch’s t-test), from 1 up to 100,000.
- Chi-square and F statistics cannot be negative.
- P-values are good to at least 9 significant digits (about 13 for most inputs; about 10 with degrees of freedom in the tens of thousands).
Worked examples
Each example is checked against the calculator on every build.
- Distribution Z, Test type Two-tailed, Test statistic 1.96 gives P-value 0.049996.
- Distribution Z, Test type Left-tailed, Test statistic -2.33 gives P-value 0.009903.Source: NIST/SEMATECH e-Handbook 1.3.6.7.1
- Distribution t, Test type Two-tailed, Test statistic 2.228, Degrees of freedom 10 gives P-value 0.050012.Source: NIST/SEMATECH e-Handbook 1.3.6.7.2 (t 0.975, 10 = 2.228)
- Distribution t, Test type Two-tailed, Test statistic 1, Degrees of freedom 1 gives P-value 0.5.
- Distribution Chi-square, Test type Right-tailed, Test statistic 18.307, Degrees of freedom 10 gives P-value 0.050001.Source: NIST/SEMATECH e-Handbook 1.3.6.7.4 (18.307 at 0.05, 10 df)
- Distribution F, Test type Right-tailed, Test statistic 4, Degrees of freedom 2, Denominator degrees of freedom 10 gives P-value 0.052922.
- Distribution F, Test type Right-tailed, Test statistic 3.326, Degrees of freedom 5, Denominator degrees of freedom 10 gives P-value 0.049993.Source: NIST/SEMATECH e-Handbook 1.3.6.7.3 (F 0.95; 5, 10 = 3.326)
How it works
The p-value is the probability, if the null hypothesis is true, of a test statistic at least as extreme as the one you got. For a statistic s and the distribution X of the statistic under the null hypothesis:
- Left-tailed: p = P(X ≤ s).
- Right-tailed: p = P(X ≥ s).
- Two-tailed: p = 2 × min(P(X ≤ s), P(X ≥ s)), and at most 1. For the symmetric Z and t distributions this is 2 × P(X ≥ |s|). For chi-square and F, which are not symmetric, it doubles the smaller tail.
The distributions:
- Z: the standard normal distribution N(0, 1).
- t: Student's t with ν degrees of freedom.
- Chi-square: χ² with k degrees of freedom. The statistic cannot be negative.
- F: F with d1 (numerator) and d2 (denominator) degrees of freedom. The statistic cannot be negative.
The page also shows both tails, P(X ≤ s) and P(X ≥ s), which add up to 1.
The evidence label follows the usual cut-offs: below 0.001 very strong, from 0.001 to below 0.01 strong, from 0.01 to below 0.05 moderate, from 0.05 to below 0.10 weak, and 0.10 or more insufficient evidence against the null hypothesis.
Assumptions
- Each tail is computed directly, not as 1 minus the other tail, so small p-values keep their relative precision.
- P-values are good to at least 9 significant digits, and about 13 for most inputs. With degrees of freedom in the tens of thousands they are good to about 10.
- Degrees of freedom may be fractional (for example Welch's t-test), from 1 up to 100,000.
- The test statistic is between −1,000,000 and 1,000,000.
Worked examples by hand
Z = 1.96, two-tailed. The standard normal table gives P(Z ≥ 1.96) = 0.0250, so p = 2 × 0.0250 = 0.0500 (0.049996 to 6 decimals). That is moderate evidence, just below the 0.05 cut-off.
Z = −2.33, left-tailed. P(Z ≤ −2.33) = 0.0099.
t = 1, 1 degree of freedom, two-tailed. With 1 degree of freedom t is the Cauchy distribution: P(|T| ≤ t) = (2 ÷ π) arctan(t). So p = 1 − (2 ÷ π) arctan(1) = 1 − (2 ÷ π)(π ÷ 4) = 0.5.
t = 2.228, 10 degrees of freedom, two-tailed. 2.228 is the t table value for 0.025 in each tail, so p = 0.0500. The old page showed 0.030184 here, which is wrong.
Chi-square = 18.307, 10 degrees of freedom, right-tailed. For an even k, P(X ≥ x) = e^(−x/2) × (1 + (x/2) + (x/2)² ÷ 2! + … + (x/2)^(k/2 − 1) ÷ (k/2 − 1)!). With x/2 = 9.1535 and 5 terms, p = 0.0500, the table value.
F = 4, d1 = 2, d2 = 10, right-tailed. For d1 = 2, P(F ≥ x) = (1 + 2x ÷ d2)^(−d2/2) = (1 + 0.8)^(−5) = 1 ÷ 18.896 = 0.0529.
Other questions people ask
What is a p-value?
A p-value is the probability of obtaining test results at least as extreme as the observed results, assuming that the null hypothesis is true. It measures the strength of evidence against the null hypothesis. A smaller p-value indicates stronger evidence against the null hypothesis.
How do I interpret p-values?
P-values are typically compared to a significance level (α): • p < 0.001: Very strong evidence against null hypothesis • p < 0.01: Strong evidence against null hypothesis • p < 0.05: Moderate evidence against null hypothesis • p < 0.10: Weak evidence against null hypothesis • p ≥ 0.10: Insufficient evidence to reject null hypothesis
What are the different types of hypothesis tests?
There are three main types of hypothesis tests: • **Left-tailed test**: Tests if a parameter is less than a specified value (H₁: parameter < value) • **Right-tailed test**: Tests if a parameter is greater than a specified value (H₁: parameter > value) • **Two-tailed test**: Tests if a parameter is different from a specified value (H₁: parameter ≠ value)
What distributions does this calculator support?
This calculator supports four common statistical distributions: • **Z-score**: Standard normal distribution N(0,1) • **T-score**: Student's t-distribution (requires degrees of freedom) • **Chi-square score**: Chi-square distribution (requires degrees of freedom) • **F-score**: F-distribution (requires two degrees of freedom)
When do I need degrees of freedom?
Degrees of freedom are required for: • **T-distribution**: Always required (typically n-1 for sample mean tests) • **Chi-square distribution**: Always required (depends on test type) • **F-distribution**: Always requires two degrees of freedom • **Z-distribution**: Not required (standard normal distribution)
What is the difference between p-value and significance level?
The p-value is calculated from your data and represents the probability of observing your results under the null hypothesis. The significance level (α) is a threshold you choose before conducting the test (commonly 0.05, 0.01, or 0.10). You reject the null hypothesis if p-value ≤ α.
Can p-values be negative?
No, p-values cannot be negative. P-values are probabilities and must be between 0 and 1. A p-value of 0 would indicate that the observed result is impossible under the null hypothesis, while a p-value of 1 would indicate that the observed result is exactly what would be expected under the null hypothesis.
What does it mean if p-value = 0.05?
A p-value of exactly 0.05 means that there is a 5% chance of observing results at least as extreme as yours if the null hypothesis is true. At the α = 0.05 significance level, you would reject the null hypothesis. However, p = 0.05 is considered borderline significant and should be interpreted with caution.
How do I choose the right significance level?
The choice of significance level depends on your field and the consequences of errors: • **α = 0.01**: Very strict, used when false positives are costly • **α = 0.05**: Standard in most fields, balances Type I and Type II errors • **α = 0.10**: More lenient, used when false negatives are costly Consider the practical implications of your decision when choosing α.
What are Type I and Type II errors?
In hypothesis testing: • **Type I error (α)**: Rejecting the null hypothesis when it's true (false positive) • **Type II error (β)**: Failing to reject the null hypothesis when it's false (false negative) The significance level α controls the probability of Type I error. The power of the test (1-β) is the probability of correctly rejecting a false null hypothesis.
When should I use a one-tailed vs two-tailed test?
Use a **one-tailed test** when you have a specific directional hypothesis (e.g., 'the new drug is better than the old one'). Use a **two-tailed test** when you're testing for any difference (e.g., 'the new drug is different from the old one'). Two-tailed tests are more conservative and are often preferred unless you have strong theoretical reasons for a directional hypothesis.
How accurate are the p-value calculations?
This calculator computes each tail of the cumulative distribution function from the incomplete gamma and beta functions, to at least 9 significant digits (about 13 for most inputs; with degrees of freedom in the tens of thousands the incomplete gamma and beta functions lose a few digits, down to about 10). It agrees with statistical software packages like R, SPSS, or SAS, and with printed tables to the places they show.
What if my test statistic is very large or very small?
Very large or small test statistics typically result in very small p-values, indicating strong evidence against the null hypothesis. The calculator works out the small tail directly, so a p-value such as 0.0000001 keeps its precision. Below about 1e-300 it shows 0.