📘 CodingMarble Learn

Similarity and Clustering (K-Means)

Clustering means putting similar items into groups without being told the groups first. Similarity is measured by distance: a smaller distance means more alike. K-means picks K centres, gives each dot to its nearest centre, moves each centre to the middle (mean) of its dots, and repeats until nothing changes.

🎬 Step-by-step story

  1. Here are dots on a table. Each dot is one student: how long they play and how long they study. Which dots belong together?
  2. We put down 3 centres (the cubes) as a first guess. The dots have no colour yet.
  3. Take one dot. A line goes to each centre. The shortest line wins, and the dot takes that centre's colour.
  4. Every dot does the same. Now each dot has the colour of its nearest centre. We have 3 groups.
  5. The centres slide to the middle of their own dots. Then the dots choose again, and a few change colour.
  6. Free play. Change K, press Next step, and press New start to see if the groups stay the same.

Tip: drag the 3D scene to turn it. Use two fingers to zoom.

🤔 Common doubts, cleared

How does the computer know which dots are alike?

It measures the distance between dots. Dots that are close are alike. Watch the lines in the 3D.

Why do the centres move?

The first centres are only a guess. Moving each one to the middle of its dots puts it where its group really is.

Why does k-means ever stop?

Each round the groups get tighter. Finally no dot wants to change, and then nothing moves any more.

Who decides how many groups there are?

You do, by choosing K. Slide K in the 3D and see what each choice gives.

Why are the dots grey at first?

Nobody has told us the groups. The colours come only after the nearest-centre step.

What is clustering?

Clustering means putting similar things in the same group. Nobody gives us the group names first. The computer finds the groups by looking at the data. This is called unsupervised learning (learning without answers given).

Each item is a dot. Its place on the table comes from its numbers, called features. In the 3D, one feature is sport hours and the other is study hours.

Similarity means distance

How do we say two dots are alike? We measure how far apart they are. Small distance = very alike. Big distance = different.

For two features, the straight-line distance is:

d = √[(x₂ − x₁)² + (y₂ − y₁)²]

This is the same Pythagoras rule you know from triangles. Example: from (0, 0) to (3, 4) the distance is √(9 + 16) = 5.

If features have very different sizes (like age and salary), shrink them to the same range first, or the big numbers will win every time.

K-means step by step

  1. Choose K, the number of groups you want.
  2. Place K centres (a first guess, often K random dots).
  3. Assign: every dot joins its nearest centre.
  4. Move: each centre goes to the mean (average position) of its dots. This new centre is called the centroid.
  5. Repeat steps 3 and 4 until no dot changes group.

Each round makes the groups tighter, so the process always stops.

Choosing K and knowing the limits

K-means does not know the right K. Try K = 2, 3, 4 and see which grouping makes sense for the real problem. Too small a K mixes different things. Too big a K cuts one real group into pieces.

A different first guess can give different groups, so people run it a few times. K-means works best when groups are round and similar in size. It is also pulled by odd dots (outliers) far from everyone.

Try it

In the 3D: before you press Next step, guess which dots will change colour. Then check. Press New start three times: do you always get the same three groups?

At home: write 10 friends' heights and shoe sizes on paper as dots. Draw 2 circles around the groups you see. Find the middle of each circle and see which dots are nearer the other middle.

Key formulas and definitions

Worked examples

1. Find the distance between the dots (1, 2) and (4, 6).

Step 1: differences are 4 − 1 = 3 and 6 − 2 = 4. Step 2: d = √(3² + 4²) = √(9 + 16) = √25 = 5.

2. A dot at (2, 2) sits between centre A at (0, 0) and centre B at (5, 5). Which centre does it join?

Distance to A = √(4 + 4) = √8 ≈ 2.8. Distance to B = √(9 + 9) = √18 ≈ 4.2. A is nearer, so the dot joins A.

3. A group has the dots (2, 4), (4, 8) and (6, 6). Where does its centre move?

Mean of x = (2 + 4 + 6) / 3 = 4. Mean of y = (4 + 8 + 6) / 3 = 6. The new centre is (4, 6).

Common mistakes

Practice quiz

1. Clustering is used to:
2. In k-means, what does K mean?
3. A smaller distance between two dots means they are:
4. The centroid of a group is its:
5. When do we stop k-means?

Practice: answer these yourself

Type or choose your answer, then press Check. Use a hint if you are stuck; the full solution appears after you answer.

Frequently asked questions

What is k-means clustering in simple words?

It is a way to split dots into K groups. Dots join the nearest centre, the centres move to the middle of their dots, and this repeats until nothing changes.

Is clustering supervised or unsupervised?

Unsupervised. The computer is not given the right groups. It finds them from the data.

How is the distance between two points found?

Use d = √[(x₂ − x₁)² + (y₂ − y₁)²]. This is the Pythagoras rule on the two differences.

Learn first

Learn next

Related lessons

All Computer Science lessons