📘 CodingMarble Learn

Data Cleaning: Outliers, Missing Values and Normalisation

Real data is messy. Before we analyse it we clean it: find and fix outliers (strange values), deal with missing values (fill or remove them), and normalise (rescale) numbers so that different measurements can be compared fairly. This is called data preprocessing.

🎬 Step-by-step story

  1. Heights of 12 students in cm. Look at the bars. One bar is huge (1650 cm) and one place is empty. The red line is the mean (average): it says 296 cm. No student is that tall!
  2. That huge bar is an outlier: a value far from the others. Here it is a typing mistake: someone typed 1650 instead of 165.
  3. We fix the typo: 1650 becomes 165. The bar drops and the mean falls to 161 cm, which makes sense.
  4. One student has no height recorded: a missing value. We fill the gap with the median (the middle value), 162 cm. Now the data is complete.
  5. Normalising rescales the numbers to run from 0 (smallest) to 1 (largest). The bars keep their order but the scale changes, so heights and weights can be compared fairly.
  6. Your turn. Press the three buttons in any order: fix the outlier, fill the gap, normalise. Watch the mean line and the message.

Tip: drag the 3D scene to turn it. Use two fingers to zoom.

🤔 Common doubts, cleared

Why does one wrong value change the mean so much?

The mean adds every value, so one huge value drags the total up. Watch the red line jump to 296 cm.

How do I know a value is an outlier and not real?

Compare it with the others and with real-world sense. 1650 cm is impossible, so it is an error; look for its cause.

Why not fill the gap with zero?

Zero would be a fake, very small height and would pull the mean down. The median is a safe, typical value.

Why normalise at all?

So features with big numbers (like income) do not overpower features with small numbers (like height).

Does normalising change the order of the data?

No. The bars stay in the same order; only the scale changes.

Why clean data?

Real data comes from people, sensors and forms. It has typing mistakes, broken sensors, blank boxes and different units. If we do not clean it, the mean, graphs and AI models all go wrong. A computer rule is simple: garbage in, garbage out.

The steps of data preprocessing: look at the data → fix wrong entries → handle missing values → handle outliers → rescale → use.

Outliers

An outlier is a value that is very far from the rest. It can be a mistake (typo, broken sensor), or a real but rare value (a very tall person, a record-breaking day).

How to spot them:

What to do: first find out why. A clear mistake: fix it (if you know the right value) or remove it. A real value: keep it, and use the median instead of the mean, because the median is hardly moved by outliers. Never delete data silently; write down what you did.

Missing values

A missing value is a blank: the child was absent, the sensor failed, the form was left empty.

Percentage missing = missing ÷ total × 100. If very many are missing, filling can mislead.

Normalisation: putting numbers on one scale

Suppose height is in cm (150 to 170) and income is in rupees (10,000 to 90,000). A computer method that compares distances would think income matters thousands of times more, just because its numbers are bigger. Normalisation rescales every feature to a common range.

Min–max normalisation (scales to 0–1): x′ = (x − min) ÷ (max − min).

Standardisation (z-score): z = (x − mean) ÷ standard deviation. Values are centred at 0.

Normalise after fixing outliers: an outlier would squash all the other values near 0.

Putting it together: a cleaning checklist

  1. Keep the raw copy safe. Work on a copy.
  2. Look: summary numbers and graphs.
  3. Fix wrong types, units and typos.
  4. Deal with missing values and write how.
  5. Check for outliers (graph + IQR) and decide with a reason.
  6. Normalise if methods need it.
  7. Record every change in a note.

Try it: clean your class data

Ask 10 classmates for their height in cm and write them in a list. Secretly add one silly value (like 1650) and leave one blank. Swap lists with a friend. Your friend must: find the outlier, decide how to fix it, fill the gap with the median, then calculate the mean before and after. Then press the buttons in the 3D above and compare.

Key formulas and definitions

Worked examples

1. Marks of 5 students: 12, 15, 14, 13, 90 (the 90 should be 19). Find the mean with the typo and the median.

Mean = (12 + 15 + 14 + 13 + 90) ÷ 5 = 144 ÷ 5 = 28.8, which is nonsense. Sorted: 12, 13, 14, 15, 90, so the median = 14. The median is hardly moved by the outlier.

2. Use the IQR rule on 2, 4, 5, 6, 7, 8, 9, 50. Is 50 an outlier?

Lower half: 2, 4, 5, 6 so Q1 = 4.5. Upper half: 7, 8, 9, 50 so Q3 = 8.5. IQR = 4. Upper fence = 8.5 + 1.5 × 4 = 14.5. Since 50 > 14.5, it is an outlier.

3. Fill the missing value with the mean of the others: 10, 12, ?, 14, 16.

Known values: 10 + 12 + 14 + 16 = 52 and there are 4 of them. Mean = 52 ÷ 4 = 13. So the missing value becomes 13.

4. Normalise the value 30 using min–max when the data run from 20 to 50.

x′ = (30 − 20) ÷ (50 − 20) = 10 ÷ 30 = 0.33 (to 2 decimals).

5. A class has mean 60 and standard deviation 10 in a test. What is the z-score for a mark of 75? Is it unusual?

z = (75 − 60) ÷ 10 = 1.5. It is above average but not unusual (|z| is below 3).

6. A student has height 160 cm (range 150 to 170) and weight 55 kg (range 40 to 70). Normalise both. What do you notice?

Height: (160 − 150) ÷ 20 = 0.5. Weight: (55 − 40) ÷ 30 = 0.5. Both are 0.5, exactly in the middle of their ranges, so we can compare them fairly.

Common mistakes

Practice quiz

1. An outlier is:
2. Which is least affected by an outlier?
3. Min–max normalisation scales data to:
4. A good way to fill a missing value when outliers exist:
5. If x = 40, min = 20, max = 60, then normalised x is:

Practice: answer these yourself

Type or choose your answer, then press Check. Use a hint if you are stuck; the full solution appears after you answer.

Frequently asked questions

What is data cleaning in simple words?

It means fixing the mistakes and gaps in data before using it: correct wrong values, fill or remove blanks, deal with odd values, and rescale numbers.

Should I always remove outliers?

No. Remove or fix them only if they are mistakes. If they are real, keep them and use methods like the median that are not easily moved by them.

What is the difference between normalisation and standardisation?

Normalisation (min–max) squeezes values into 0 to 1. Standardisation (z-score) centres values at 0 using the mean and standard deviation. Both put features on a common scale.

Learn first

Learn next

Related lessons

All Maths lessons