Why clean data?
Real data comes from people, sensors and forms. It has typing mistakes, broken sensors, blank boxes and different units. If we do not clean it, the mean, graphs and AI models all go wrong. A computer rule is simple: garbage in, garbage out.
The steps of data preprocessing: look at the data → fix wrong entries → handle missing values → handle outliers → rescale → use.
Outliers
An outlier is a value that is very far from the rest. It can be a mistake (typo, broken sensor), or a real but rare value (a very tall person, a record-breaking day).
How to spot them:
- Draw a graph (bar, dot plot, box plot). Odd values stand out.
- IQR rule: find Q1 and Q3, IQR = Q3 − Q1. Anything below Q1 − 1.5 × IQR or above Q3 + 1.5 × IQR is flagged.
- z-score: z = (x − mean) ÷ standard deviation. Many people flag |z| above 3.
What to do: first find out why. A clear mistake: fix it (if you know the right value) or remove it. A real value: keep it, and use the median instead of the mean, because the median is hardly moved by outliers. Never delete data silently; write down what you did.
Missing values
A missing value is a blank: the child was absent, the sensor failed, the form was left empty.
- Remove the row (fine if only a few are missing).
- Fill (impute) with the mean (for fairly even data), the median (when there are outliers) or the mode (most common value, for categories).
- Mark that it was missing, so that you remember it was filled.
- Best of all: go back and collect it again.
Percentage missing = missing ÷ total × 100. If very many are missing, filling can mislead.
Normalisation: putting numbers on one scale
Suppose height is in cm (150 to 170) and income is in rupees (10,000 to 90,000). A computer method that compares distances would think income matters thousands of times more, just because its numbers are bigger. Normalisation rescales every feature to a common range.
Min–max normalisation (scales to 0–1): x′ = (x − min) ÷ (max − min).
Standardisation (z-score): z = (x − mean) ÷ standard deviation. Values are centred at 0.
Normalise after fixing outliers: an outlier would squash all the other values near 0.
Putting it together: a cleaning checklist
- Keep the raw copy safe. Work on a copy.
- Look: summary numbers and graphs.
- Fix wrong types, units and typos.
- Deal with missing values and write how.
- Check for outliers (graph + IQR) and decide with a reason.
- Normalise if methods need it.
- Record every change in a note.
Try it: clean your class data
Ask 10 classmates for their height in cm and write them in a list. Secretly add one silly value (like 1650) and leave one blank. Swap lists with a friend. Your friend must: find the outlier, decide how to fix it, fill the gap with the median, then calculate the mean before and after. Then press the buttons in the 3D above and compare.
Key formulas and definitions
- Mean = sum of values ÷ number of values; median = middle value after sorting
- IQR = Q3 − Q1; outlier if x < Q1 − 1.5 × IQR or x > Q3 + 1.5 × IQR
- z = (x − mean) ÷ standard deviation (flag |z| > 3)
- Min–max: x′ = (x − min) ÷ (max − min) (range 0 to 1)
- Percentage missing = missing ÷ total × 100
Worked examples
1. Marks of 5 students: 12, 15, 14, 13, 90 (the 90 should be 19). Find the mean with the typo and the median.
Mean = (12 + 15 + 14 + 13 + 90) ÷ 5 = 144 ÷ 5 = 28.8, which is nonsense. Sorted: 12, 13, 14, 15, 90, so the median = 14. The median is hardly moved by the outlier.
2. Use the IQR rule on 2, 4, 5, 6, 7, 8, 9, 50. Is 50 an outlier?
Lower half: 2, 4, 5, 6 so Q1 = 4.5. Upper half: 7, 8, 9, 50 so Q3 = 8.5. IQR = 4. Upper fence = 8.5 + 1.5 × 4 = 14.5. Since 50 > 14.5, it is an outlier.
3. Fill the missing value with the mean of the others: 10, 12, ?, 14, 16.
Known values: 10 + 12 + 14 + 16 = 52 and there are 4 of them. Mean = 52 ÷ 4 = 13. So the missing value becomes 13.
4. Normalise the value 30 using min–max when the data run from 20 to 50.
x′ = (30 − 20) ÷ (50 − 20) = 10 ÷ 30 = 0.33 (to 2 decimals).
5. A class has mean 60 and standard deviation 10 in a test. What is the z-score for a mark of 75? Is it unusual?
z = (75 − 60) ÷ 10 = 1.5. It is above average but not unusual (|z| is below 3).
6. A student has height 160 cm (range 150 to 170) and weight 55 kg (range 40 to 70). Normalise both. What do you notice?
Height: (160 − 150) ÷ 20 = 0.5. Weight: (55 − 40) ÷ 30 = 0.5. Both are 0.5, exactly in the middle of their ranges, so we can compare them fairly.
Common mistakes
- Deleting an outlier without asking why it happened. It may be real and important.
- Filling missing values with the mean when there are big outliers. Use the median then.
- Normalising before fixing outliers. One outlier squashes all other values near 0.
- Not keeping the raw data and not noting what was changed.