📘 CodingMarble Learn

Model Evaluation: How Good Is a Model?

A model is a simple copy of the real world that makes predictions. To evaluate it, we compare its predictions with the real answers. In machine learning we split data into a training set (to learn) and a test set (to check on new data). A confusion matrix counts true positives (TP), false positives (FP), false negatives (FN) and true negatives (TN). Accuracy = (TP + TN) ÷ total. Precision = TP ÷ (TP + FP). Recall = TP ÷ (TP + FN). F1 = 2PR ÷ (P + R). Science models (like the inverse-square law or the particle model) are judged the same way: do predictions match measurements, and where does the model stop working?

🎬 Step-by-step story

  1. A model makes predictions. We check them against the real answers.
  2. Split the data. Train on most of it. Test on the hidden part.
  3. A confusion matrix sorts results into four boxes: TP, FP, FN, TN.
  4. Accuracy = right answers ÷ all answers. Here it is 85%.
  5. Precision: how many alarms were real? Recall: how many real ones were caught?
  6. Free play: move the threshold. Watch precision and recall pull against each other.

Tip: drag the 3D scene to turn it. Use two fingers to zoom.

🤔 Common doubts, cleared

Why not test on the training data? It is easier.

The model has already seen it, so it can score high just by memorising. Step 2 keeps the gold test blocks apart.

What is the difference between FP and FN?

FP is a false alarm (said yes, was no). FN is a miss (said no, was yes). See the orange and purple stacks in step 3.

If accuracy is 99%, is the model always good?

No. If 99% of cases are "no", always saying "no" gets 99% but finds nothing. Step 4 explains this trap.

Precision and recall look alike. How do I remember?

Both have TP on top. Precision divides by the top row (all predicted yes); recall divides by the left column (all really yes). Step 5 greys out the other boxes.

Can I get 100% precision and 100% recall together?

Only with a perfect model. Usually raising one lowers the other. Move the threshold in step 6 and watch.

How are physics models like ML models?

Both predict, and both are judged by comparing predictions with real measurements, as in step 1.

What is a model, and why evaluate it?

A model is a simplified version of something real. It helps us explain or predict.

No model is perfect. Evaluating a model means checking how close its predictions are to reality, and finding where it fails. A good model is simple, gives correct predictions, and we know its limits.

Train-test split

We divide the data into two parts:

Why? A model can memorise the training data and still fail on new data. This is called overfitting. A model that is too simple and fails even on training data is underfitting. Testing on unseen data is like an exam with new questions.

The confusion matrix

For a yes/no model (spam or not), every test result falls in one of four boxes:

Rows are usually the prediction and columns the reality (some books swap them, so always read the labels).

Accuracy, precision, recall and F1

Accuracy = (TP + TN) ÷ total. Share of all answers that were right.

Precision = TP ÷ (TP + FP). Of everything the model called yes, how much was really yes?

Recall = TP ÷ (TP + FN). Of everything that was really yes, how much did the model find?

F1 score = 2 × P × R ÷ (P + R). One number that balances precision and recall.

Which one matters?

Accuracy misleads when classes are unbalanced (99 healthy, 1 sick). For a disease test, a miss is dangerous, so recall matters most. For a spam filter, hiding a real email is annoying, so precision matters. Changing the threshold trades one for the other.

Judging scientific models

Physics models are tested the same way: predict, measure, compare.

In technology (solar panels, MRI, phones), recognise which laws are used and where the simple model stops working.

Try it

Make your own "rain model": each evening predict "rain" or "no rain" tomorrow from the clouds. After 10 days, fill a confusion matrix and work out accuracy, precision and recall. Then use the 3D slider to see how a stricter rule changes the scores.

Key formulas and definitions

Worked examples

1. A test set has 200 photos. A cat-detector gets 170 right. What is its accuracy?

Accuracy = 170 ÷ 200 = 0.85 = 85%.

2. A spam filter gives TP = 40, FP = 10, FN = 5, TN = 45. Find accuracy, precision, recall and F1.

Total = 100. Accuracy = (40 + 45) ÷ 100 = 85%. Precision = 40 ÷ 50 = 0.80. Recall = 40 ÷ 45 ≈ 0.889. F1 = 2 × 0.80 × 0.889 ÷ (0.80 + 0.889) = 1.422 ÷ 1.689 ≈ 0.84.

3. Out of 1,000 people, 10 have a rare disease. A model always says "healthy". Find accuracy and recall. Is it a good model?

TN = 990, FN = 10, TP = 0, FP = 0. Accuracy = 990 ÷ 1000 = 99%. Recall = 0 ÷ 10 = 0%. It is useless: it never finds a sick person. High accuracy can hide a bad model when classes are unbalanced.

4. A lamp gives intensity 36 W/m² at 1 m. Using the inverse-square model, what is it at 3 m?

I₂ = 36 × (1/3)² = 36 ÷ 9 = 4 W/m². If measurement gives about 4, the model fits; near a big reflector it may not.

Common mistakes

Practice quiz

1. Which data is used to check a model at the end?
2. Predicted spam, but it was a normal email. This is a:
3. Accuracy =
4. For a cancer screening test, which score matters most?
5. If distance from a point light doubles, the intensity becomes:

Practice: answer these yourself

Type or choose your answer, then press Check. Use a hint if you are stuck; the full solution appears after you answer.

Frequently asked questions

What is a confusion matrix?

A 2×2 table that counts true positives, false positives, false negatives and true negatives, so you can see exactly what kind of mistakes a model makes.

What is a good F1 score?

Closer to 1 is better. What counts as "good" depends on the task; compare models on the same test data.

Why do all models have limits?

A model leaves out details to stay simple. When those details matter (very small, very fast, very far), its predictions stop matching reality.

Where this is taught

NetherlandsHAVO 5 (eindexamenjaar)Physics and technology
NetherlandsVWO 6 (eindexamenjaar)Laws of nature and models
CBSE (India)Class 10Part B: Evaluating Models

Learn first

Related lessons

All Computer Science lessons