What is a model, and why evaluate it?
A model is a simplified version of something real. It helps us explain or predict.
- Physical model: a globe, a model of the atom.
- Mathematical model: a formula, like distance = speed × time.
- Computer / AI model: a program that learned from data, like a spam filter.
No model is perfect. Evaluating a model means checking how close its predictions are to reality, and finding where it fails. A good model is simple, gives correct predictions, and we know its limits.
Train-test split
We divide the data into two parts:
- Training set (often 70-80%): the model learns from it.
- Test set (20-30%): kept hidden, used only at the end.
Why? A model can memorise the training data and still fail on new data. This is called overfitting. A model that is too simple and fails even on training data is underfitting. Testing on unseen data is like an exam with new questions.
The confusion matrix
For a yes/no model (spam or not), every test result falls in one of four boxes:
- TP (true positive): predicted yes, really yes.
- FP (false positive): predicted yes, really no. A false alarm.
- FN (false negative): predicted no, really yes. A miss.
- TN (true negative): predicted no, really no.
Rows are usually the prediction and columns the reality (some books swap them, so always read the labels).
Accuracy, precision, recall and F1
Accuracy = (TP + TN) ÷ total. Share of all answers that were right.
Precision = TP ÷ (TP + FP). Of everything the model called yes, how much was really yes?
Recall = TP ÷ (TP + FN). Of everything that was really yes, how much did the model find?
F1 score = 2 × P × R ÷ (P + R). One number that balances precision and recall.
Which one matters?
Accuracy misleads when classes are unbalanced (99 healthy, 1 sick). For a disease test, a miss is dangerous, so recall matters most. For a spam filter, hiding a real email is annoying, so precision matters. Changing the threshold trades one for the other.
Judging scientific models
Physics models are tested the same way: predict, measure, compare.
- Universality and scale: a good law works from tiny to huge (gravity for an apple and for the Moon), but some models only work at one scale (Newton's mechanics fails near the speed of light or inside atoms).
- Order of magnitude: a quick estimate to the nearest power of 10 checks if an answer is sensible.
- Inverse-square law: light, sound and gravity spread over a sphere of area 4πr², so intensity ∝ 1/r². Double the distance, one quarter the intensity.
- Conservation laws (energy, momentum, charge): any model that breaks them is wrong.
- Analogies: electric current is like water flow; useful, but every analogy breaks down somewhere.
In technology (solar panels, MRI, phones), recognise which laws are used and where the simple model stops working.
Try it
Make your own "rain model": each evening predict "rain" or "no rain" tomorrow from the clouds. After 10 days, fill a confusion matrix and work out accuracy, precision and recall. Then use the 3D slider to see how a stricter rule changes the scores.
Key formulas and definitions
- Accuracy = (TP + TN) ÷ (TP + TN + FP + FN)
- Precision = TP ÷ (TP + FP)
- Recall (sensitivity) = TP ÷ (TP + FN)
- F1 = 2 × Precision × Recall ÷ (Precision + Recall)
- Error rate = 1 − accuracy
- Inverse-square law: I ∝ 1/r², so I₂ = I₁ × (r₁/r₂)²
Worked examples
1. A test set has 200 photos. A cat-detector gets 170 right. What is its accuracy?
Accuracy = 170 ÷ 200 = 0.85 = 85%.
2. A spam filter gives TP = 40, FP = 10, FN = 5, TN = 45. Find accuracy, precision, recall and F1.
Total = 100. Accuracy = (40 + 45) ÷ 100 = 85%. Precision = 40 ÷ 50 = 0.80. Recall = 40 ÷ 45 ≈ 0.889. F1 = 2 × 0.80 × 0.889 ÷ (0.80 + 0.889) = 1.422 ÷ 1.689 ≈ 0.84.
3. Out of 1,000 people, 10 have a rare disease. A model always says "healthy". Find accuracy and recall. Is it a good model?
TN = 990, FN = 10, TP = 0, FP = 0. Accuracy = 990 ÷ 1000 = 99%. Recall = 0 ÷ 10 = 0%. It is useless: it never finds a sick person. High accuracy can hide a bad model when classes are unbalanced.
4. A lamp gives intensity 36 W/m² at 1 m. Using the inverse-square model, what is it at 3 m?
I₂ = 36 × (1/3)² = 36 ÷ 9 = 4 W/m². If measurement gives about 4, the model fits; near a big reflector it may not.
Common mistakes
- Testing a model on the same data it was trained on. The score looks great but means little.
- Trusting accuracy alone when one class is rare.
- Mixing up precision and recall: precision divides by everything predicted yes; recall divides by everything really yes.
- Thinking a model that fits well is "true". Every model has limits; say where it stops working.