📘 CodingMarble Learn

Machine Learning

Machine learning (ML) is a way for computers to learn a rule from examples instead of being told the rule. We give data made of features (inputs) and, in supervised learning, labels (answers). The machine fits a model: regression predicts a number, a decision tree asks yes/no questions to pick a class, and k-means clustering groups unlabelled data. We train on most of the data, test on data it never saw, and measure accuracy or error. A model that only memorises (overfits) fails on new data.

🎬 Step-by-step story

  1. Each dot is one fruit. Its size and sweetness are the features. Its colour shows the label: apple or lemon.
  2. Regression: the green line starts in a bad place. It slowly turns until the red error sticks are as short as possible. Now it can predict a rent.
  3. Decision tree: first question, sweetness more than 5? Then, size more than 6 cm? Each part of the plane gives one answer.
  4. K-means: the dots have no labels. Two centres drop in. Each dot joins the nearest centre, the centres move to the middle, and we repeat.
  5. Testing: grey dots were hidden during learning. Green ring = the model got it right, red ring = wrong. 4 out of 5 gives 80% accuracy.
  6. Your turn: choose k from 1 to 5 and watch the dots sort themselves into groups.

Tip: drag the 3D scene to turn it. Use two fingers to zoom.

🤔 Common doubts, cleared

Is a feature the same as a label?

No. Features are inputs (size, sweetness). The label is the answer (apple or lemon). In the 3D, position = features, colour = label.

How does the machine know which line is best?

It measures the errors (red sticks) and keeps changing the slope and intercept to make the total squared error smaller.

Who decides the questions in a decision tree?

The training algorithm tries many questions and keeps the one that separates the classes best at each split.

How can it group data when there are no answers?

It uses distance: close points are probably similar. The centres keep moving until the groups stop changing.

Why not test on the training data?

The model has already seen it. Only hidden data shows real performance, as the green and red rings show.

How do I choose k in k-means?

Try several values and see where adding a group no longer helps much. Try k = 1 to 5 in free play.

What is machine learning?

In normal programming, a person writes the rule: if sweetness > 5 then apple. In machine learning, we give the computer many examples and it finds the rule itself.

ML is one part of artificial intelligence (AI). Humans learn from a few examples and use common sense. Machines need many examples, but they are very fast, never get tired and can look at millions of numbers.

To define an ML problem, ask: What do I want to predict? Do I have past examples with answers? How will I check if the machine is good?

Data, features and labels

Each row of data is one example. The columns we use as inputs are features (size, sweetness, area of a flat). The answer we want is the label (apple/lemon, rent).

Preparing data: remove mistakes and missing values, put numbers on similar scales, and turn words into numbers. Bad data gives a bad model: “garbage in, garbage out”.

Three common models: regression, decision trees, k-means

Linear regression (predict a number)

Fit a line y = mx + c through the data. The error is the gap between each real point and the line. Training changes m and c to make the mean squared error as small as possible.

Decision tree (pick a class)

A tree of yes/no questions. Each question splits the data. The ends, called leaves, give the answer. Trees are easy for people to read.

K-means clustering (find groups)

Choose k. Place k centres. Repeat: each point joins its nearest centre; each centre moves to the average of its points. Stop when nothing changes.

Choosing a model: number to predict → regression; category with labels → classification (tree, neural network); no labels → clustering.

Training, testing and evaluation

Split the data: about 80% for training, 20% for testing. The test data stays hidden while learning, like an exam paper you have not seen.

ML can also be unfair if the data is biased. Check whose data is missing, and keep personal data private.

Try it: design an AI solution

Pick a small problem, such as guessing if it will rain tomorrow. 1) Define: rain yes/no. 2) Data: write down today's cloud cover, humidity and wind for 20 days, and whether it rained next day. 3) Model: make a two-question tree by hand. 4) Test it on the last 5 days. Count how many it got right. In the 3D, change k and see how the groups change.

Key formulas and definitions

Worked examples

1. A spam filter is tested on 200 emails and gets 184 right. Find its accuracy.

Accuracy = 184 ÷ 200 × 100% = 92%.

2. A regression model for flat rent is y = 0.5x + 3 (y in thousand rupees, x = area in m²÷10). Predict rent for a 60 m² flat.

x = 60 ÷ 10 = 6. y = 0.5 × 6 + 3 = 6. Predicted rent ≈ ₹6,000.

3. Points 2, 4, 10, 12 on a line, k = 2, starting centres 2 and 4. Do one round of k-means.

Nearest centre: 2 → 2; 4 → 4; 10 → 4; 12 → 4. Groups {2} and {4, 10, 12}. New centres: 2 and (4 + 10 + 12)/3 ≈ 8.67. Next round: 4 is nearer 2 (distance 2) than 8.67, so groups become {2, 4} and {10, 12}, centres 3 and 11. Then nothing changes.

4. A model scores 99% on training data but 60% on test data. What is happening?

It is overfitting: it memorised the training examples instead of learning the general pattern. Fix: more data, a simpler model, or stop training earlier.

Common mistakes

Practice quiz

1. In supervised learning, the data has:
2. Which model predicts a number like a house price?
3. K-means is an example of:
4. Why keep test data hidden during training?
5. A model that is perfect on training data but poor on new data is:

Practice: answer these yourself

Type or choose your answer, then press Check. Use a hint if you are stuck; the full solution appears after you answer.

Frequently asked questions

What is machine learning in simple words?

It is teaching a computer with examples so it can find a pattern and make predictions on new data.

What are the main types of machine learning?

Supervised (with labels), unsupervised (no labels, e.g. clustering) and reinforcement (learning from rewards).

What is the difference between AI and ML?

AI is the big goal of making machines act smart. ML is one method of AI where machines learn from data.

Where this is taught

NetherlandsHAVO 5 (eindexamenjaar)Elective theme: Cognitive computing
NetherlandsVWO 6 (eindexamenjaar)Elective theme: Cognitive computing
CBSE (India)Class 10Part B: Advanced Concepts of Modeling in AI
CBSE (India)Class 11Machine Learning Algorithms
South Korea중학교 2학년Artificial intelligence
South Korea고등학교 2학년AI and learning
South Korea고등학교 2학년AI project
South Korea고등학교 2학년Text data
South Korea고등학교 2학년Image data
South Korea고등학교 2학년Artificial intelligence
South Korea고등학교 2학년Science and logic
South Korea고등학교 3학년Classification and prediction
Germany (Bavaria)Jahrgangsstufe 11Artificial intelligence
China九年级(初三)Module: AI and smart society
China高三Sel.4 Introduction to AI

Learn first

Learn next

Related lessons

All Computer Science lessons