acalculator

What is my linear regression line?

Enter your x values and matching y values to get the line of best fit and how well it fits.

Your numbers

Read as: 1; 2; 3; 4; 5
Read as: 2; 4; 5; 4; 5
Line of best fit
y = 0.6x + 2.2

The line of best fit through 5 points is y = 0.6x + 2.2.

Slope (b)
0.6
Intercept (a)
2.2
Correlation (r)
0.774597
r²
0.6
Residual standard deviation (s)
0.894427
Predicted y
5.8
Number of points (n)
5

Line of best fit: y = 0.6x + 2.2. The line of best fit through 5 points is y = 0.6x + 2.2.

The line of best fit

How to calculate

Fits the least-squares line y = a + bx to paired x and y values and reports the slope, intercept, correlation r, r², residual standard deviation, and a prediction.

Example with the default inputs (x values [1, 2, 3, 4, 5], y values [2, 4, 5, 4, 5], New x value 6): The line of best fit through 5 points is y = 0.6x + 2.2.

Method: b = Σ(x − x̄)(y − ȳ) ÷ Σ(x − x̄)²; a = ȳ − b x̄; the line is y = a + bx.

  • The first x value goes with the first y value, and so on, so both lists must be the same length.
  • The line minimises the sum of squared vertical distances from the points (ordinary least squares).
  • r and r² are not defined when every y is the same, and the residual standard deviation needs at least 3 points.

Machine-readable copies: Markdown, JSON.

Worked examples

Each example is checked against the calculator on every build.

  1. x values 1, 2, 3, 4, 5, y values 2, 4, 5, 4, 5, New x value 6 gives Line of best fit y = 0.6x + 2.2, Slope (b) 0.6, Intercept (a) 2.2, Correlation (r) 0.774597, r² 0.6, Residual standard deviation (s) 0.894427, Predicted y 5.8.Source: hand calculation in content.mdx: Sxx = 10, Sxy = 6, Syy = 6, b = 0.6, a = 4 − 0.6 × 3 = 2.2, r = 6 ÷ √60, s = √(2.4 ÷ 3); Python statistics module, linear_regression and correlation
  2. x values 0.2, 337.4, 118.2, 884.6, 10.1, 226.5, 666.3, 996.3, 448.6, 777, 558.2, 0.4, 0.6, 775.5, 666.9, 338, 447.5, 11.6, 556, 228.1, 995.8, 887.6, 120.2, 0.3, 0.3, 556.8, 339.1, 887.2, 999, 779, 11.1, 118.3, 229.2, 669.1, 448.9, 0.5, y values 0.1, 338.8, 118.1, 888, 9.2, 228.1, 668.5, 998.5, 449.1, 778.9, 559.2, 0.3, 0.1, 778.1, 668.8, 339.3, 448.9, 10.8, 557.7, 228.3, 998, 888.8, 119.6, 0.3, 0.6, 557.6, 339.3, 888, 998.5, 778.9, 10.2, 117.6, 228.9, 668.4, 449.2, 0.2 gives Slope (b) 1.002117, Intercept (a) -0.262323, Residual standard deviation (s) 0.884796, r² 0.999994.Source: NIST StRD linear least squares dataset Norris: certified B1, B0, residual standard deviation, and R-squared (15 digits, so tolerance 1e-12; Python agrees to 4.7e-14)
  3. x values 0, 1, 2, 3, y values 10, 7, 4, 1 gives Line of best fit y = −3x + 10, Slope (b) -3, Intercept (a) 10, Correlation (r) -1, r² 1, Residual standard deviation (s) 0.Source: hand calculation in content.mdx: every step in x lowers y by 3, and y = 10 at x = 0

How it works

For n pairs (x₁, y₁), …, (xₙ, yₙ) with means x̄ and ȳ:

  • Sxx = Σ(xᵢ − x̄)², Syy = Σ(yᵢ − ȳ)², Sxy = Σ(xᵢ − x̄)(yᵢ − ȳ).
  • Slope: b = Sxy ÷ Sxx (NIST e-Handbook 4.4.3.1).
  • Intercept: a = ȳ − b × x̄.
  • Line of best fit: y = a + bx, shown as y = bx + a, or y = bx − |a| when a is below 0 before rounding. b and |a| are rounded to at most 4 decimal places, half away from zero, working on the shortest decimal form of the number (as JavaScript's Intl.NumberFormat does); trailing zeros are dropped, thousands get commas, a negative b gets a minus sign (−), and a value that rounds to 0 shows as 0.
  • Correlation: r = Sxy ÷ √(Sxx × Syy), and r² = r².
  • Residual standard deviation: s = √(SSE ÷ (n − 2)), where SSE = Σ(yᵢ − a − bxᵢ)², the sum of squared residuals (NIST e-Handbook 4.4.3.1, with p = 2 parameters).
  • Predicted y at the new x value you enter: a + b × x. Leave that field empty to skip it.

The chart draws the fitted line from the smallest to the largest x value (stretched to the new x value when it is outside them) and marks the prediction.

Assumptions

  • The lists are matched by position: the first x goes with the first y. Both lists must have the same number of values, at least 2; otherwise there is no answer.
  • At least two x values must differ; if all x values are the same, there is no answer.
  • When every y value is the same, the line is flat and r and r² are not defined, so they are not shown.
  • The residual standard deviation needs at least 3 points and is not shown for 2.
  • This is ordinary least squares with an intercept: it minimises the vertical distances only, and treats x as known exactly.

Worked examples by hand

x = 1, 2, 3, 4, 5 and y = 2, 4, 5, 4, 5 (the default). The means are x̄ = 15 ÷ 5 = 3 and ȳ = 20 ÷ 5 = 4. The differences from the means are x − x̄ = −2, −1, 0, 1, 2 and y − ȳ = −2, 0, 1, 0, 1. So Sxx = 4 + 1 + 0 + 1 + 4 = 10, Sxy = 4 + 0 + 0 + 0 + 2 = 6, and Syy = 4 + 0 + 1 + 0 + 1 = 6.

  • b = 6 ÷ 10 = 0.6, and a = 4 − 0.6 × 3 = 2.2, so the line is y = 0.6x + 2.2.
  • r = 6 ÷ √(10 × 6) = 6 ÷ √60 = 0.7746, and r² = 36 ÷ 60 = 0.6.
  • The fitted values are 2.8, 3.4, 4.0, 4.6, 5.2, and the residuals are −0.8, 0.6, 1.0, −0.6, −0.2. Their squares add up to SSE = 0.64 + 0.36 + 1 + 0.36 + 0.04 = 2.4, so s = √(2.4 ÷ 3) = √0.8 = 0.8944.
  • At x = 6, the line predicts y = 2.2 + 0.6 × 6 = 5.8.

x = 0, 1, 2, 3 and y = 10, 7, 4, 1. Each step of 1 in x lowers y by 3, so every point is on the line y = −3x + 10: b = −3, a = 10, r = −1, r² = 1, and every residual is 0, so s = 0.

NIST's Norris data (36 points from an ozone monitor calibration). The calculator gives b = 1.00211681802045, a = −0.262323073774029, s = 0.884796396144373, and r² = 0.999993745883712, matching NIST's certified values to 13 significant figures or more.

Other questions people ask

How do I calculate a linear regression line by hand?

Find the means x̄ and ȳ. Work out Sxx = Σ(x − x̄)² and Sxy = Σ(x − x̄)(y − ȳ). The slope is b = Sxy ÷ Sxx and the intercept is a = ȳ − b × x̄. The line is y = a + bx. For x = 1 to 5 and y = 2, 4, 5, 4, 5, this gives b = 6 ÷ 10 = 0.6 and a = 4 − 0.6 × 3 = 2.2.

What does “least squares” mean?

For any straight line, each point has a vertical distance to it, called a residual. The least-squares line is the one line that makes the sum of the squared residuals as small as possible. Squaring counts points above and below the line alike and weighs large misses more.

What do r and r² tell me?

r, the correlation coefficient, runs from −1 to 1. Its sign matches the slope, and values near ±1 mean the points lie close to a straight line. r² is the share of the variation in y that the line explains: r² = 0.6 means the line accounts for 60% of it.

What is the residual standard deviation?

It is the typical vertical distance of a point from the line, in the units of y: √(SSE ÷ (n − 2)), where SSE is the sum of squared residuals. It divides by n − 2 because the line used up two numbers, the slope and the intercept. NIST calls it the residual standard deviation.

Can I use the line to predict y for a new x?

Yes: enter the x value and read y = a + bx. Predictions are most reliable for x values inside the range of your data. Far outside that range, there is no evidence that the straight-line pattern still holds.

Does a strong correlation mean x causes y?

No. A line that fits well shows that x and y move together in your data. Both could be driven by a third factor, or the link could run the other way. Showing cause needs a designed experiment or other evidence.

Why does the calculator say my x values must not all be the same?

If every x is the same, the points lie on a vertical line and Sxx = 0, so the slope Sxy ÷ Sxx has no value. Least squares needs at least two different x values.