# What is my linear regression line?

Fits the least-squares line y = a + bx to paired x and y values and reports the slope, intercept, correlation r, r², residual standard deviation, and a prediction.

- Page: https://www.acalculator.org/statistics/linear-regression-calculator
- JSON spec: https://www.acalculator.org/statistics/linear-regression-calculator.json
- Version: 750cd52555bb

## Default answer

Example with the default inputs (x values [1, 2, 3, 4, 5], y values [2, 4, 5, 4, 5], New x value 6): The line of best fit through 5 points is y = 0.6x + 2.2.

## Inputs

| Key | Label | Description |
| --- | --- | --- |
| x | x values | The x values (the predictor), in order, separated by commas, spaces, semicolons, or new lines. |
| y | y values | The y values (the response), one for each x value, in the same order. |
| at | New x value | An x value at which to read y from the fitted line. Leave empty to skip. |

## Outputs

| Key | Label | Description |
| --- | --- | --- |
| equation | Line of best fit | The least-squares line, with the slope and intercept rounded to 4 decimal places. |
| slope | Slope (b) | How much y changes, on average, when x goes up by 1: Sxy ÷ Sxx. |
| intercept | Intercept (a) | The value of y on the line where x = 0. |
| r | Correlation (r) | Pearson’s correlation coefficient, from −1 to 1. Not shown when every y is the same. |
| r2 | r² | The share of the variation in y that the line explains, from 0 to 1. Not shown when every y is the same. |
| residualSd | Residual standard deviation (s) | The typical distance of a point from the line: √(SSE ÷ (n − 2)). Needs 3 or more points. |
| predicted | Predicted y | The value of y on the line at the new x value: a + b × x. |
| n | Number of points (n) | How many (x, y) pairs there are. |

## Method

b = Σ(x − x̄)(y − ȳ) ÷ Σ(x − x̄)²; a = ȳ − b x̄; the line is y = a + bx.

## Assumptions

- The first x value goes with the first y value, and so on, so both lists must be the same length.
- The line minimises the sum of squared vertical distances from the points (ordinary least squares).
- r and r² are not defined when every y is the same, and the residual standard deviation needs at least 3 points.

## Worked examples

1. x = 1 or 2, y = 2 or 4, at = 6 gives equation = y = 0.6x + 2.2, slope = 0.6, intercept = 2.2, r = 0.774597, r2 = 0.6, residualSd = 0.894427, predicted = 5.8. Source: hand calculation in content.mdx: Sxx = 10, Sxy = 6, Syy = 6, b = 0.6, a = 4 − 0.6 × 3 = 2.2, r = 6 ÷ √60, s = √(2.4 ÷ 3); Python statistics module, linear_regression and correlation.
2. x = 0.2 or 337.4, y = 0.1 or 338.8 gives slope = 1.002117, intercept = -0.262323, residualSd = 0.884796, r2 = 0.999994. Source: NIST StRD linear least squares dataset Norris: certified B1, B0, residual standard deviation, and R-squared (15 digits, so tolerance 1e-12; Python agrees to 4.7e-14).
3. x = 0 or 1, y = 10 or 7 gives equation = y = −3x + 10, slope = -3, intercept = 10, r = -1, r2 = 1, residualSd = 0. Source: hand calculation in content.mdx: every step in x lowers y by 3, and y = 10 at x = 0.

## FAQ

### How do I calculate a linear regression line by hand?

Find the means x̄ and ȳ. Work out Sxx = Σ(x − x̄)² and Sxy = Σ(x − x̄)(y − ȳ). The slope is b = Sxy ÷ Sxx and the intercept is a = ȳ − b × x̄. The line is y = a + bx. For x = 1 to 5 and y = 2, 4, 5, 4, 5, this gives b = 6 ÷ 10 = 0.6 and a = 4 − 0.6 × 3 = 2.2.

### What does “least squares” mean?

For any straight line, each point has a vertical distance to it, called a residual. The least-squares line is the one line that makes the sum of the squared residuals as small as possible. Squaring counts points above and below the line alike and weighs large misses more.

### What do r and r² tell me?

r, the correlation coefficient, runs from −1 to 1. Its sign matches the slope, and values near ±1 mean the points lie close to a straight line. r² is the share of the variation in y that the line explains: r² = 0.6 means the line accounts for 60% of it.

### What is the residual standard deviation?

It is the typical vertical distance of a point from the line, in the units of y: √(SSE ÷ (n − 2)), where SSE is the sum of squared residuals. It divides by n − 2 because the line used up two numbers, the slope and the intercept. NIST calls it the residual standard deviation.

### Can I use the line to predict y for a new x?

Yes: enter the x value and read y = a + bx. Predictions are most reliable for x values inside the range of your data. Far outside that range, there is no evidence that the straight-line pattern still holds.

### Does a strong correlation mean x causes y?

No. A line that fits well shows that x and y move together in your data. Both could be driven by a third factor, or the link could run the other way. Showing cause needs a designed experiment or other evidence.

### Why does the calculator say my x values must not all be the same?

If every x is the same, the points lie on a vertical line and Sxx = 0, so the slope Sxy ÷ Sxx has no value. Least squares needs at least two different x values.

## Sources

- NIST/SEMATECH e-Handbook of Statistical Methods, section 4.4.3.1, Least Squares (slope, intercept, and residual standard deviation). https://www.itl.nist.gov/div898/handbook/pmd/section4/pmd431.htm
- NIST/SEMATECH e-Handbook of Statistical Methods, section 4.1.4.1, Linear Least Squares Regression. https://www.itl.nist.gov/div898/handbook/pmd/section1/pmd141.htm
- NIST Statistical Reference Datasets, linear least squares regression: Norris, certified values. https://www.itl.nist.gov/div898/strd/lls/data/Norris.shtml
