Linear Regression Calculator

Paste or type your x values and y values, one list each in the same order. The calculator fits the best straight line and tells you how well it fits.

Separate numbers with commas, spaces or new lines.
Separate numbers with commas, spaces or new lines. The nth y goes with the nth x.
The fitted line is evaluated at this x.

Regression line

ŷ = 47.6071 + 4.6429x

Slope b

4.6429

Intercept a

47.6071

Correlation r

0.9974

r² (coefficient of determination)

0.9949

Standard error of the estimate

0.8797

Predicted y

94.0357

Points used

8

How it works

Linear regression finds the one straight line that comes closest to a set of points, where "closest" means the sum of the squared vertical distances from the points to the line is as small as possible. That line is ŷ = a + bx: b is the slope (how much y changes for each unit of x) and a is the intercept (the value of y where x = 0).

The slope comes from how x and y vary together, Σ(x − x̄)(y − ȳ), divided by how x varies on its own, Σ(x − x̄)². The intercept is then fixed by the fact that the line always passes through the point of means (x̄, ȳ).

r² says what fraction of the spread in y the line accounts for: 1 is a perfect fit, 0 means x tells you nothing about y. The standard error of the estimate is the typical size of a residual (a point's vertical distance from the line), in the units of y, so it tells you how far a prediction is likely to be off.

Formula

x̄ = Σx ÷ n          ȳ = Σy ÷ n
Sxx = Σ(x − x̄)²     Syy = Σ(y − ȳ)²     Sxy = Σ(x − x̄)(y − ȳ)
b  = Sxy ÷ Sxx                     (slope)
a  = ȳ − b·x̄                       (intercept)
r² = Sxy² ÷ (Sxx · Syy)            r = ±√r², with the sign of b
SE = √( Σ(y − ŷ)² ÷ (n − 2) )      (standard error of the estimate)
ŷ  = a + b·x                       (prediction)

Example

Eight weeks of practice (x = 1 to 8) and test scores y = 52, 57, 61, 68, 70, 75, 80, 85. The means are x̄ = 4.5 and ȳ = 68.5; Sxy = 195 and Sxx = 42, so the slope is 195 ÷ 42 = 4.6429 points per week and the intercept is 68.5 − 4.6429 × 4.5 = 47.6071. The line is ŷ = 47.6071 + 4.6429x.

Syy = 910, so r² = 195² ÷ (42 × 910) = 0.9949 and r = 0.9974: the line explains 99.5% of the variation in scores. The standard error of the estimate is 0.8797 points, and the predicted score after 10 weeks is 47.6071 + 4.6429 × 10 = 94.0357.

Assumptions and limitations

  • This is ordinary least squares: x is taken as exact and only the vertical distance from each point to the line is minimised. If both variables carry measurement error of similar size, a different fit is appropriate.
  • A high r² means the line fits these points well. It does not mean x causes y, and it says nothing about what happens outside the range of x you entered.
  • The standard error divides by n − 2, so at least three points are needed.
  • The x and y lists are paired by position: the third y belongs to the third x. A missing value shifts every later pair, so check the two lists are the same length and in the same order.

Frequently asked questions

What is the difference between r and r²?

r is the correlation coefficient, between −1 and 1, whose sign says whether y rises or falls with x and whose size says how tightly the points hug a line. r² is its square, between 0 and 1, and has the direct meaning "fraction of the variation in y explained by the line". r = 0.7 sounds strong but explains only 49% of the variation.

Why does the result show — for r when all my y values are the same?

If every y is equal there is no variation in y to explain, so the formula for r divides zero by zero. The line itself is still well defined: it is horizontal at that y value, with slope 0.