📘 CodingMarble Learn

Linear Regression and the Least Squares Line

Linear regression finds the straight line ŷ = a + bx that best follows paired data (x, y). A residual is the gap between a real point and the line: e = y − ŷ. The least squares line makes the sum of squared residuals as small as possible. Its slope is b = Sxy ÷ Sxx and it always passes through the mean point (x̄, ȳ). We use it to predict y from x, but only inside the data range, and a strong link does not prove that x causes y.

🎬 Step-by-step story

  1. Each blue dot is one student: hours studied (x) and marks (y). The dots go up from left to right.
  2. Let us guess a line. This flat green line just says "everyone gets the average mark, 56".
  3. Red sticks show the gap from each dot to the line. This gap is called the residual: real y minus line y.
  4. Square each gap. The orange squares show the squared residuals. Their total area tells how bad the line is.
  5. Now the line tilts until the total square area is as small as it can be. This is the least squares line: ŷ = 29.4 + 5.9x.
  6. Your turn: move a and b, try to beat the best line, and slide x to predict a mark. Tap "Best fit" to check.

Tip: drag the 3D scene to turn it. Use two fingers to zoom.

🤔 Common doubts, cleared

Why do we square the residuals instead of just adding them?

Points above the line give positive gaps and points below give negative ones, so they cancel. Squaring makes every gap positive, and big misses count much more.

Why is a flat line through the average not good enough?

It ignores x completely. In the 3D the squares under the flat line are huge; tilting the line shrinks them a lot.

Why does the best line go through (x̄, ȳ)?

Because a = ȳ − b·x̄. Put x = x̄ in the line and you get ŷ = ȳ.

Is a residual measured straight down or at a right angle to the line?

Straight up or down (vertically), because we are predicting y. The red sticks in the 3D are vertical.

Can I trust a prediction for 20 hours of study?

No. Our data only go from 1 to 8 hours. Going far outside is extrapolation and the line may give silly answers. Slide x in the free-play step to see.

What if the points do not go in a straight line?

Then a straight line is a poor model, and the residuals show a pattern. Use a curve instead.

What is a regression model?

We often have paired data: two numbers for each person or thing, like (hours studied, marks). We draw them as dots on a scatter diagram.

If the dots lie roughly along a straight line, we can describe them with a linear regression model:

ŷ = a + bx

Real points do not sit exactly on the line. The full model is y = a + bx + e, where e is a random error.

Residuals: how far each point is from the line

A residual is the vertical gap between a real point and the line:

e = y − ŷ

Positive residual: the point is above the line. Negative residual: the point is below it. For the least squares line, the residuals always add up to 0.

A residual plot (x against e) helps us check the model. If the residuals are scattered randomly around 0, a straight line fits well. If they make a curve or a funnel, a straight line is not a good model.

The least squares method

Some residuals are positive and some negative, so we cannot just add them. We square them first. The least squares line is the line that makes the sum of squared residuals, Σe², as small as possible.

Formulas (n pairs of data):

Because a = ȳ − b·x̄, the line always passes through the mean point (x̄, ȳ). The slope b has the same sign as the correlation coefficient r. In fact b = r × (sy ÷ sx).

Line of y on x and line of x on y

The line above predicts y from x (y on x). To predict x from y you need a different line, x = c + dy with d = Sxy ÷ Syy. The two lines cross at (x̄, ȳ) and are the same only when r = ±1.

Using the line to predict, and its limits

Put a value of x into ŷ = a + bx to get a prediction.

Correlation is not causation

A strong straight-line link does not prove that x causes y. Ice-cream sales and drowning cases both rise in summer. Hot weather drives both. This hidden cause is called a lurking (or confounding) variable. Only a fair experiment can show cause and effect.

The coefficient of determination r² tells what fraction of the variation in y the line explains. r = 0.9 gives r² = 0.81, so the line explains 81% of the variation.

Try it: your own regression line

Ask 6 friends or family members for their height (cm) and arm span (cm). Plot the points. Find x̄, ȳ, Sxx and Sxy, then the line. Is the slope close to 1? Use the 3D: move a and b to make the orange squares as small as you can, then press "Best fit" to see how close you were.

What exams ask

Typical questions: find the regression line from a table (4–6 marks); interpret a and b in words; predict a value and say whether it is reliable; find residuals; choose y on x or x on y; explain why correlation does not mean causation. Always state units and write the equation with ŷ.

Key formulas and definitions

Worked examples

1. The line is ŷ = 29.4 + 5.9x (x = hours studied, y = marks). Predict the mark for 5 hours.

ŷ = 29.4 + 5.9 × 5 = 29.4 + 29.5 = 58.9 ≈ 59 marks.

2. For the same line, a student who studied 4 hours scored 50. Find the residual.

ŷ = 29.4 + 5.9 × 4 = 53.0. Residual e = y − ŷ = 50 − 53.0 = −3.0. The point is 3 marks below the line.

3. Explain what the slope 5.9 and intercept 29.4 mean.

Slope: each extra hour of study goes with about 5.9 more marks on average. Intercept: a student who studies 0 hours is predicted to get about 29 marks (only a rough idea, since x = 0 may be outside the data).

4. Data: x = 1, 2, 3, 4, 5 and y = 3, 5, 6, 8, 8. Find the least squares line of y on x.

n = 5, x̄ = 3, ȳ = 30 ÷ 5 = 6. Σx² = 55, so Sxx = 55 − 5 × 9 = 10. Σxy = 3 + 10 + 18 + 32 + 40 = 103, so Sxy = 103 − 5 × 3 × 6 = 13. b = 13 ÷ 10 = 1.3. a = 6 − 1.3 × 3 = 2.1. Line: ŷ = 2.1 + 1.3x.

5. Using ŷ = 2.1 + 1.3x from the last example, find all residuals and check their sum.

ŷ = 3.4, 4.7, 6.0, 7.3, 8.6. Residuals: 3 − 3.4 = −0.4, 5 − 4.7 = 0.3, 0, 0.7, −0.6. Sum = −0.4 + 0.3 + 0 + 0.7 − 0.6 = 0. Squared sum = 0.16 + 0.09 + 0 + 0.49 + 0.36 = 1.1.

6. r = 0.8, the standard deviation of x is 2 and of y is 5, x̄ = 10, ȳ = 40. Find the line of y on x and predict y at x = 12.

b = r × sy ÷ sx = 0.8 × 5 ÷ 2 = 2. a = 40 − 2 × 10 = 20. ŷ = 20 + 2x. At x = 12: ŷ = 44.

7. A line ŷ = 12 + 0.5x was made from data with x between 20 and 60. Is a prediction at x = 150 reliable?

No. x = 150 is far outside 20–60, so this is extrapolation. The straight-line pattern may not continue, so the prediction ŷ = 87 should not be trusted.

Common mistakes

Practice quiz

1. A residual is:
2. The least squares line makes which quantity smallest?
3. The least squares line always passes through:
4. If r is negative, the slope b is:
5. Predicting outside the data range is called:

Practice: answer these yourself

Type or choose your answer, then press Check. Use a hint if you are stuck; the full solution appears after you answer.

Frequently asked questions

What is linear regression in simple words?

It is a way to draw the best straight line through scattered data points so you can predict one quantity from another.

What is the difference between correlation and regression?

Correlation (r) measures how strong and in which direction a straight-line link is. Regression gives the actual equation of the line, ŷ = a + bx, to make predictions.

Why is it called "least squares"?

Because the chosen line makes the sum of the squares of the residuals the least (smallest) possible.

Where this is taught

CBSE (India)Class 11Descriptive Statistics
South Korea고등학교 2학년Prediction and optimisation
South Korea고등학교 2학년Modelling and evaluation
South Korea고등학교 3학년Classification and prediction
South Korea고등학교 3학년Optimisation
FranceTerminaleStudy themes
FranceTerminaleMathematics (2019 programme)
China高三Ch.8 Paired data

Learn first

Learn next

Related lessons

All Maths lessons