What is machine learning?
In normal programming, a person writes the rule: if sweetness > 5 then apple. In machine learning, we give the computer many examples and it finds the rule itself.
ML is one part of artificial intelligence (AI). Humans learn from a few examples and use common sense. Machines need many examples, but they are very fast, never get tired and can look at millions of numbers.
To define an ML problem, ask: What do I want to predict? Do I have past examples with answers? How will I check if the machine is good?
Data, features and labels
Each row of data is one example. The columns we use as inputs are features (size, sweetness, area of a flat). The answer we want is the label (apple/lemon, rent).
- Supervised learning: examples come with labels. The model learns input → answer.
- Unsupervised learning: no labels. The model finds groups or patterns.
- Reinforcement learning: the model tries actions and gets rewards, like a game player.
Preparing data: remove mistakes and missing values, put numbers on similar scales, and turn words into numbers. Bad data gives a bad model: “garbage in, garbage out”.
Three common models: regression, decision trees, k-means
Linear regression (predict a number)
Fit a line y = mx + c through the data. The error is the gap between each real point and the line. Training changes m and c to make the mean squared error as small as possible.
Decision tree (pick a class)
A tree of yes/no questions. Each question splits the data. The ends, called leaves, give the answer. Trees are easy for people to read.
K-means clustering (find groups)
Choose k. Place k centres. Repeat: each point joins its nearest centre; each centre moves to the average of its points. Stop when nothing changes.
Choosing a model: number to predict → regression; category with labels → classification (tree, neural network); no labels → clustering.
Training, testing and evaluation
Split the data: about 80% for training, 20% for testing. The test data stays hidden while learning, like an exam paper you have not seen.
- Accuracy = correct predictions ÷ total predictions.
- For regression we use the average error.
- Overfitting: the model memorises training data, even its noise, and does badly on new data.
- Underfitting: the model is too simple to catch the pattern.
ML can also be unfair if the data is biased. Check whose data is missing, and keep personal data private.
Try it: design an AI solution
Pick a small problem, such as guessing if it will rain tomorrow. 1) Define: rain yes/no. 2) Data: write down today's cloud cover, humidity and wind for 20 days, and whether it rained next day. 3) Model: make a two-question tree by hand. 4) Test it on the last 5 days. Count how many it got right. In the 3D, change k and see how the groups change.
Key formulas and definitions
- Feature = an input column; label = the answer column
- Linear regression: y = mx + c
- Mean squared error = average of (actual − predicted)²
- Accuracy = correct predictions ÷ total predictions × 100%
- K-means: assign to nearest centre → move centre to mean → repeat
- Usual split: 80% training data, 20% test data
Worked examples
1. A spam filter is tested on 200 emails and gets 184 right. Find its accuracy.
Accuracy = 184 ÷ 200 × 100% = 92%.
2. A regression model for flat rent is y = 0.5x + 3 (y in thousand rupees, x = area in m²÷10). Predict rent for a 60 m² flat.
x = 60 ÷ 10 = 6. y = 0.5 × 6 + 3 = 6. Predicted rent ≈ ₹6,000.
3. Points 2, 4, 10, 12 on a line, k = 2, starting centres 2 and 4. Do one round of k-means.
Nearest centre: 2 → 2; 4 → 4; 10 → 4; 12 → 4. Groups {2} and {4, 10, 12}. New centres: 2 and (4 + 10 + 12)/3 ≈ 8.67. Next round: 4 is nearer 2 (distance 2) than 8.67, so groups become {2, 4} and {10, 12}, centres 3 and 11. Then nothing changes.
4. A model scores 99% on training data but 60% on test data. What is happening?
It is overfitting: it memorised the training examples instead of learning the general pattern. Fix: more data, a simpler model, or stop training earlier.
Common mistakes
- Testing the model on the same data used for training — it looks great but tells you nothing.
- Thinking the machine “understands” like a human. It only finds patterns in numbers.
- Using regression to predict a category or clustering when labels are available.
- Ignoring biased or dirty data. A model can only be as good as its data.