What are we trying to find?
Many problems ask for the smallest value of something: the lowest cost, the smallest error. This smallest value is the minimum of a function.
Picture the function as a hill. The height at any spot is the value of the function (here, the error, also called loss). We want the lowest spot.
Slope tells us the way down
The slope (or gradient) tells us how steep the ground is and which way is uphill. In maths it is the derivative. For f(x) = x², the slope at x is 2x.
If the slope is positive, going left lowers the error. If it is negative, going right lowers it. So we always move against the slope. That is why it is called descent. For a hill with two directions, the gradient is the pair of slopes, one for each direction.
The update rule and the learning rate
Every step uses one rule:
new x = old x − learning rate × slope
The learning rate (often written η or α) is a small number you choose. It sets how big each step is. A steep slope gives a big step; near the bottom the slope is small, so steps shrink by themselves.
- Too big: the ball jumps over the minimum and may bounce further away.
- Too small: it is safe but needs very many steps.
- Just right: it goes steadily down to the minimum.
When do we stop?
We stop when the slope is almost zero (the ball is nearly flat), when the error stops getting smaller, or after a fixed number of steps. Bumpy hills can have several valleys. The ball may stop in a small valley that is not the lowest one. Starting from different places helps.
This is the idea of optimisation: finding the best value. See also the lesson on optimisation.
Try it yourself
In the 3D: set the learning rate to 0.1, press Run and count how many steps it needs. Now set it to 0.5. Then set it to about 1.2 and watch the ball zigzag.
At home: close your eyes in a room with a gentle slope, such as a ramp or sloping ground. Take small steps in the direction your feet say is downhill. Try big jumps and see how you overshoot.
Key formulas and definitions
- new x = old x − learning rate × slope
- x₍ₙ₊₁₎ = xₙ − η · f'(xₙ)
- For f(x) = x²: slope = 2x
- Stop when |slope| is almost 0 or the error stops falling
Worked examples
1. f(x) = x². Find the slope at x = 4.
Slope = 2x = 2 × 4 = 8. It is positive, so going left lowers the error.
2. Start at x = 4 with learning rate 0.1 on f(x) = x². Find the new x after one step.
Slope = 8. new x = 4 − 0.1 × 8 = 4 − 0.8 = 3.2.
3. Start at x = 2 with learning rate 0.25 on f(x) = x². Find x after two steps.
Step 1: slope = 4, x = 2 − 0.25 × 4 = 1. Step 2: slope = 2, x = 1 − 0.25 × 2 = 0.5. The ball is getting close to the minimum at 0.
4. Start at x = 3 with learning rate 0.5 on f(x) = x². What happens?
Slope = 6. new x = 3 − 0.5 × 6 = 0. One step lands exactly on the minimum. This is a very good learning rate for this function.
5. Start at x = 3 with learning rate 1 on f(x) = x². What happens?
Step 1: slope = 6, x = 3 − 6 = −3. Step 2: slope = −6, x = −3 + 6 = 3. The ball jumps between 3 and −3 forever. The learning rate is too big.
6. f(x) = (x − 5)². Start at x = 1 with learning rate 0.25. Find x after one step.
Slope = 2(x − 5) = 2(1 − 5) = −8. new x = 1 − 0.25 × (−8) = 1 + 2 = 3. The slope is negative, so x moves right, towards 5.
Common mistakes
- Moving along the slope instead of against it. Descent means we go downhill, so we subtract the slope.
- Thinking a bigger learning rate is always better. Too big can overshoot or even blow up.
- Using a learning rate of 0 or a very tiny one and wondering why nothing seems to move.
- Thinking gradient descent always finds the lowest point of the whole hill. It can stop in a small valley.