📘 CodingMarble Learn

Gradient Descent: Finding the Lowest Point

Gradient descent finds the minimum of a function by taking small steps downhill. At each step, new x = old x − learning rate × slope. A learning rate that is too big overshoots; too small is very slow. It is how machine learning models reduce their error.

🎬 Step-by-step story

  1. This is a hill of error. The lower the ground, the fewer mistakes the computer makes. The green ring is the lowest point.
  2. A red ball starts high up. The blue arrow shows the steepest way downhill. The ball cannot see the bottom, it only feels the slope under it.
  3. Now we take steps. Each step moves the ball a little along the arrow. The yellow dots show its path and the error number gets smaller.
  4. The learning rate is the size of a step. Make it too big and the ball jumps over the valley and zigzags from side to side.
  5. Make it too small and the ball moves, but very slowly. Many steps are needed to get down.
  6. Your turn. Change the learning rate, press One step or Run, and try to reach the green ring in the fewest steps.

Tip: drag the 3D scene to turn it. Use two fingers to zoom.

🤔 Common doubts, cleared

What does the height of the hill mean?

It is the error. The higher the ground, the more mistakes. The green ring is the place with the fewest mistakes.

Why does the arrow point downhill?

The slope tells which way is uphill. The arrow points the opposite way, so the ball always goes down.

Why do the steps get smaller near the bottom?

A step is the learning rate times the slope. Near the bottom the ground is almost flat, so the slope is small and so is the step. Watch the yellow dots bunch together.

Why does the ball jump from side to side?

Each step is so long that the ball lands on the other side of the valley, even higher up on one side. Then it jumps back. The learning rate is too big.

Is a small learning rate better, then?

It is safe, but needs very many steps. The best choice is in the middle: fast and still steady.

What are we trying to find?

Many problems ask for the smallest value of something: the lowest cost, the smallest error. This smallest value is the minimum of a function.

Picture the function as a hill. The height at any spot is the value of the function (here, the error, also called loss). We want the lowest spot.

Slope tells us the way down

The slope (or gradient) tells us how steep the ground is and which way is uphill. In maths it is the derivative. For f(x) = x², the slope at x is 2x.

If the slope is positive, going left lowers the error. If it is negative, going right lowers it. So we always move against the slope. That is why it is called descent. For a hill with two directions, the gradient is the pair of slopes, one for each direction.

The update rule and the learning rate

Every step uses one rule:

new x = old x − learning rate × slope

The learning rate (often written η or α) is a small number you choose. It sets how big each step is. A steep slope gives a big step; near the bottom the slope is small, so steps shrink by themselves.

When do we stop?

We stop when the slope is almost zero (the ball is nearly flat), when the error stops getting smaller, or after a fixed number of steps. Bumpy hills can have several valleys. The ball may stop in a small valley that is not the lowest one. Starting from different places helps.

This is the idea of optimisation: finding the best value. See also the lesson on optimisation.

Try it yourself

In the 3D: set the learning rate to 0.1, press Run and count how many steps it needs. Now set it to 0.5. Then set it to about 1.2 and watch the ball zigzag.

At home: close your eyes in a room with a gentle slope, such as a ramp or sloping ground. Take small steps in the direction your feet say is downhill. Try big jumps and see how you overshoot.

Key formulas and definitions

Worked examples

1. f(x) = x². Find the slope at x = 4.

Slope = 2x = 2 × 4 = 8. It is positive, so going left lowers the error.

2. Start at x = 4 with learning rate 0.1 on f(x) = x². Find the new x after one step.

Slope = 8. new x = 4 − 0.1 × 8 = 4 − 0.8 = 3.2.

3. Start at x = 2 with learning rate 0.25 on f(x) = x². Find x after two steps.

Step 1: slope = 4, x = 2 − 0.25 × 4 = 1. Step 2: slope = 2, x = 1 − 0.25 × 2 = 0.5. The ball is getting close to the minimum at 0.

4. Start at x = 3 with learning rate 0.5 on f(x) = x². What happens?

Slope = 6. new x = 3 − 0.5 × 6 = 0. One step lands exactly on the minimum. This is a very good learning rate for this function.

5. Start at x = 3 with learning rate 1 on f(x) = x². What happens?

Step 1: slope = 6, x = 3 − 6 = −3. Step 2: slope = −6, x = −3 + 6 = 3. The ball jumps between 3 and −3 forever. The learning rate is too big.

6. f(x) = (x − 5)². Start at x = 1 with learning rate 0.25. Find x after one step.

Slope = 2(x − 5) = 2(1 − 5) = −8. new x = 1 − 0.25 × (−8) = 1 + 2 = 3. The slope is negative, so x moves right, towards 5.

Common mistakes

Practice quiz

1. Gradient descent moves the ball:
2. The update rule is:
3. If the learning rate is too big, the ball may:
4. For f(x) = x², the slope at x = 3 is:
5. A very small learning rate makes learning:

Practice: answer these yourself

Type or choose your answer, then press Check. Use a hint if you are stuck; the full solution appears after you answer.

Frequently asked questions

What is gradient descent in simple words?

It is a way to find the lowest point of a function by repeatedly checking the slope and taking a small step downhill.

What is the learning rate?

It is the number that sets the size of each step. Small means slow and safe, big means fast but risky.

Why is gradient descent used in machine learning?

A model has an error that depends on millions of numbers. Gradient descent adjusts those numbers a little at a time so the error keeps getting smaller.

Where this is taught

South Korea고등학교 2학년Prediction and optimisation
South Korea고등학교 3학년Optimisation

Learn first

Learn next

Related lessons

All Maths lessons