What is clustering?
Clustering means putting similar things in the same group. Nobody gives us the group names first. The computer finds the groups by looking at the data. This is called unsupervised learning (learning without answers given).
Each item is a dot. Its place on the table comes from its numbers, called features. In the 3D, one feature is sport hours and the other is study hours.
Similarity means distance
How do we say two dots are alike? We measure how far apart they are. Small distance = very alike. Big distance = different.
For two features, the straight-line distance is:
d = √[(x₂ − x₁)² + (y₂ − y₁)²]
This is the same Pythagoras rule you know from triangles. Example: from (0, 0) to (3, 4) the distance is √(9 + 16) = 5.
If features have very different sizes (like age and salary), shrink them to the same range first, or the big numbers will win every time.
K-means step by step
- Choose K, the number of groups you want.
- Place K centres (a first guess, often K random dots).
- Assign: every dot joins its nearest centre.
- Move: each centre goes to the mean (average position) of its dots. This new centre is called the centroid.
- Repeat steps 3 and 4 until no dot changes group.
Each round makes the groups tighter, so the process always stops.
Choosing K and knowing the limits
K-means does not know the right K. Try K = 2, 3, 4 and see which grouping makes sense for the real problem. Too small a K mixes different things. Too big a K cuts one real group into pieces.
A different first guess can give different groups, so people run it a few times. K-means works best when groups are round and similar in size. It is also pulled by odd dots (outliers) far from everyone.
Try it
In the 3D: before you press Next step, guess which dots will change colour. Then check. Press New start three times: do you always get the same three groups?
At home: write 10 friends' heights and shoe sizes on paper as dots. Draw 2 circles around the groups you see. Find the middle of each circle and see which dots are nearer the other middle.
Key formulas and definitions
- Distance: d = √[(x₂ − x₁)² + (y₂ − y₁)²]
- Centroid (mean position): x̄ = (x₁ + x₂ + … + xₙ) / n, ȳ = (y₁ + y₂ + … + yₙ) / n
- Key words: cluster = a group, centroid = middle of a group, K = number of groups, feature = a number that describes a dot
Worked examples
1. Find the distance between the dots (1, 2) and (4, 6).
Step 1: differences are 4 − 1 = 3 and 6 − 2 = 4. Step 2: d = √(3² + 4²) = √(9 + 16) = √25 = 5.
2. A dot at (2, 2) sits between centre A at (0, 0) and centre B at (5, 5). Which centre does it join?
Distance to A = √(4 + 4) = √8 ≈ 2.8. Distance to B = √(9 + 9) = √18 ≈ 4.2. A is nearer, so the dot joins A.
3. A group has the dots (2, 4), (4, 8) and (6, 6). Where does its centre move?
Mean of x = (2 + 4 + 6) / 3 = 4. Mean of y = (4 + 8 + 6) / 3 = 6. The new centre is (4, 6).
Common mistakes
- Thinking K-means finds K by itself. You must choose K.
- Forgetting to move the centres after assigning dots. Assign and move are two separate steps.
- Mixing features with very different sizes, such as height in cm and income in lakhs, without scaling them.
- Believing one run is always the best. A different first guess can give different groups.