What is a regression model?
We often have paired data: two numbers for each person or thing, like (hours studied, marks). We draw them as dots on a scatter diagram.
If the dots lie roughly along a straight line, we can describe them with a linear regression model:
ŷ = a + bx
- x is the explanatory (independent) variable: the one we know.
- ŷ ("y-hat") is the predicted value of the response variable y.
- b is the slope: how much ŷ changes when x goes up by 1.
- a is the intercept: the value of ŷ when x = 0.
Real points do not sit exactly on the line. The full model is y = a + bx + e, where e is a random error.
Residuals: how far each point is from the line
A residual is the vertical gap between a real point and the line:
e = y − ŷ
Positive residual: the point is above the line. Negative residual: the point is below it. For the least squares line, the residuals always add up to 0.
A residual plot (x against e) helps us check the model. If the residuals are scattered randomly around 0, a straight line fits well. If they make a curve or a funnel, a straight line is not a good model.
The least squares method
Some residuals are positive and some negative, so we cannot just add them. We square them first. The least squares line is the line that makes the sum of squared residuals, Σe², as small as possible.
Formulas (n pairs of data):
- x̄ = Σx ÷ n, ȳ = Σy ÷ n
- Sxx = Σ(x − x̄)² = Σx² − n·x̄²
- Sxy = Σ(x − x̄)(y − ȳ) = Σxy − n·x̄·ȳ
- b = Sxy ÷ Sxx
- a = ȳ − b·x̄
Because a = ȳ − b·x̄, the line always passes through the mean point (x̄, ȳ). The slope b has the same sign as the correlation coefficient r. In fact b = r × (sy ÷ sx).
Line of y on x and line of x on y
The line above predicts y from x (y on x). To predict x from y you need a different line, x = c + dy with d = Sxy ÷ Syy. The two lines cross at (x̄, ȳ) and are the same only when r = ±1.
Using the line to predict, and its limits
Put a value of x into ŷ = a + bx to get a prediction.
- Interpolation: x is inside the range of the data. Usually reliable if r is close to ±1.
- Extrapolation: x is outside the range. Risky, because the pattern may change. (Studying 30 hours would not give 200 marks!)
Correlation is not causation
A strong straight-line link does not prove that x causes y. Ice-cream sales and drowning cases both rise in summer. Hot weather drives both. This hidden cause is called a lurking (or confounding) variable. Only a fair experiment can show cause and effect.
The coefficient of determination r² tells what fraction of the variation in y the line explains. r = 0.9 gives r² = 0.81, so the line explains 81% of the variation.
Try it: your own regression line
Ask 6 friends or family members for their height (cm) and arm span (cm). Plot the points. Find x̄, ȳ, Sxx and Sxy, then the line. Is the slope close to 1? Use the 3D: move a and b to make the orange squares as small as you can, then press "Best fit" to see how close you were.
What exams ask
Typical questions: find the regression line from a table (4–6 marks); interpret a and b in words; predict a value and say whether it is reliable; find residuals; choose y on x or x on y; explain why correlation does not mean causation. Always state units and write the equation with ŷ.
Key formulas and definitions
- ŷ = a + bx
- Residual e = y − ŷ
- b = Sxy ÷ Sxx
- a = ȳ − b·x̄
- Sxx = Σx² − n·x̄²
- Sxy = Σxy − n·x̄·ȳ
- b = r · (sy ÷ sx)
- r² = share of variation in y explained by the line
Worked examples
1. The line is ŷ = 29.4 + 5.9x (x = hours studied, y = marks). Predict the mark for 5 hours.
ŷ = 29.4 + 5.9 × 5 = 29.4 + 29.5 = 58.9 ≈ 59 marks.
2. For the same line, a student who studied 4 hours scored 50. Find the residual.
ŷ = 29.4 + 5.9 × 4 = 53.0. Residual e = y − ŷ = 50 − 53.0 = −3.0. The point is 3 marks below the line.
3. Explain what the slope 5.9 and intercept 29.4 mean.
Slope: each extra hour of study goes with about 5.9 more marks on average. Intercept: a student who studies 0 hours is predicted to get about 29 marks (only a rough idea, since x = 0 may be outside the data).
4. Data: x = 1, 2, 3, 4, 5 and y = 3, 5, 6, 8, 8. Find the least squares line of y on x.
n = 5, x̄ = 3, ȳ = 30 ÷ 5 = 6. Σx² = 55, so Sxx = 55 − 5 × 9 = 10. Σxy = 3 + 10 + 18 + 32 + 40 = 103, so Sxy = 103 − 5 × 3 × 6 = 13. b = 13 ÷ 10 = 1.3. a = 6 − 1.3 × 3 = 2.1. Line: ŷ = 2.1 + 1.3x.
5. Using ŷ = 2.1 + 1.3x from the last example, find all residuals and check their sum.
ŷ = 3.4, 4.7, 6.0, 7.3, 8.6. Residuals: 3 − 3.4 = −0.4, 5 − 4.7 = 0.3, 0, 0.7, −0.6. Sum = −0.4 + 0.3 + 0 + 0.7 − 0.6 = 0. Squared sum = 0.16 + 0.09 + 0 + 0.49 + 0.36 = 1.1.
6. r = 0.8, the standard deviation of x is 2 and of y is 5, x̄ = 10, ȳ = 40. Find the line of y on x and predict y at x = 12.
b = r × sy ÷ sx = 0.8 × 5 ÷ 2 = 2. a = 40 − 2 × 10 = 20. ŷ = 20 + 2x. At x = 12: ŷ = 44.
7. A line ŷ = 12 + 0.5x was made from data with x between 20 and 60. Is a prediction at x = 150 reliable?
No. x = 150 is far outside 20–60, so this is extrapolation. The straight-line pattern may not continue, so the prediction ŷ = 87 should not be trusted.
Common mistakes
- Adding residuals to judge the fit. They always add up to 0 for the best line; you must square them.
- Mixing up the two lines: use y on x to predict y, and x on y to predict x.
- Extrapolating far outside the data range and trusting the answer.
- Saying "x causes y" just because the correlation is strong.