What is data processing and central tendency?
Data processing means working on raw data to get useful results. One common job is to find a measure of central tendency: a single value that stands for the whole set. The three measures are mean, median and mode.
Geographers use them to compare places: mean annual rainfall, median size of farms, the most common crop.
Mean: sharing equally
Ungrouped data (direct method)
Mean (x̄) = Σx ÷ N. Add all values and divide by their number.
Grouped data
- Direct method: find the midpoint (x) of each class, then x̄ = Σfx ÷ N.
- Assumed mean (indirect) method: choose a midpoint A near the centre, find d = x − A, then x̄ = A + Σfd ÷ N. This keeps numbers small.
- Step deviation: d′ = (x − A) ÷ i, then x̄ = A + (Σfd′ ÷ N) × i, where i is the class width.
The mean uses every value, so it is affected by very large or very small values.
Median: the middle value
Ungrouped data
Arrange values in ascending order. Median position = (N + 1) ÷ 2. If N is even, take the mean of the two middle values.
Grouped data
Find cumulative frequencies (cf). The median class is the first class whose cf ≥ N ÷ 2. Then
Median = l + ((N/2 − cf) ÷ f) × i
l = lower limit of the median class, cf = cumulative frequency of the class before it, f = frequency of the median class, i = class width.
The median is not pulled by extreme values.
Mode: the most frequent value
Ungrouped data
The value that appears most often. A set can be unimodal (one mode), bimodal (two) or multimodal, or have no mode.
Grouped data
The modal class has the highest frequency. Then
Mode = l + ((f1 − f0) ÷ (2f1 − f0 − f2)) × i
f1 = frequency of the modal class, f0 = of the class before, f2 = of the class after.
Mode is the only average that also works for categories (e.g. the most common soil type).
Comparing mean, median and mode
- Symmetrical (balanced) data: mean = median = mode.
- Positively skewed (a long tail of high values): mode < median < mean.
- Negatively skewed (a tail of low values): mean < median < mode.
- Empirical rule for moderately skewed data: Mode ≈ 3 Median − 2 Mean.
| Measure | Good for | Weakness |
|---|---|---|
| Mean | Uses all values; further maths | Pulled by extremes |
| Median | Skewed data (incomes, farm size) | Ignores the size of other values |
| Mode | Most common item; categories | May not exist or may be more than one |
Key formulas and definitions
- Mean (ungrouped) x̄ = Σx ÷ N
- Mean (grouped, direct) x̄ = Σfx ÷ N
- Mean (assumed mean) x̄ = A + Σfd ÷ N, d = x − A
- Median position (ungrouped) = (N + 1) ÷ 2
- Median (grouped) = l + ((N/2 − cf) ÷ f) × i
- Mode (grouped) = l + ((f1 − f0) ÷ (2f1 − f0 − f2)) × i
- Mode ≈ 3 Median − 2 Mean
Worked examples
1. Find the mean of 5, 9, 5, 13, 6, 8, 3 rainy days.
Σx = 49, N = 7. Mean = 49 ÷ 7 = 7 days.
2. Find the median of 5, 9, 5, 13, 6, 8, 3.
In order: 3, 5, 5, 6, 8, 9, 13. Position = (7 + 1) ÷ 2 = 4th. Median = 6.
3. Find the median of 12, 4, 9, 7.
In order: 4, 7, 9, 12. N = 4 (even). Middle two are 7 and 9. Median = (7 + 9) ÷ 2 = 8.
4. Find the mode of 21, 23, 21, 25, 23, 21, 30 °C.
21 comes 3 times, 23 twice, others once. Mode = 21 °C.
5. Rainfall classes 0–10, 10–20, 20–30, 30–40 have f = 3, 7, 7, 3. Find the mean by the direct method.
Midpoints x: 5, 15, 25, 35. fx: 15, 105, 175, 105. Σfx = 400, N = 20. Mean = 400 ÷ 20 = 20 cm.
6. Find the same mean by the assumed mean method with A = 25.
d = x − 25: −20, −10, 0, 10. fd: −60, −70, 0, 30. Σfd = −100. Mean = 25 + (−100 ÷ 20) = 25 − 5 = 20 cm.
7. Find the median of the same grouped data.
cf: 3, 10, 17, 20. N/2 = 10. First cf ≥ 10 is 10 (class 10–20). l = 10, cf before = 3, f = 7, i = 10. Median = 10 + ((10 − 3) ÷ 7) × 10 = 10 + 10 = 20 cm.
8. Classes 0–10, 10–20, 20–30, 30–40 have f = 2, 5, 9, 4. Find the mode.
Modal class 20–30 (f1 = 9), f0 = 5, f2 = 4, l = 20, i = 10. Mode = 20 + ((9 − 5) ÷ (18 − 5 − 4)) × 10 = 20 + (4 ÷ 9) × 10 ≈ 24.4.
Common mistakes
- Finding the median without first putting the values in order.
- Using class limits instead of midpoints when finding the mean of grouped data.
- In the median formula, using the cf of the median class itself instead of the class before it.
- Thinking the mean is always the best average. For skewed data the median is fairer.