Types of data: qualitative, quantitative, primary and secondary
Data means the facts a researcher collects. There are two big ways to sort it.
Words or numbers?
- Qualitative data is in words. Example: "I felt nervous before the test." It gives rich detail. But it is hard to compare and can be judged in a biased way.
- Quantitative data is in numbers. Example: a stress score of 7 out of 10. It is easy to count, compare and put in a graph. But it can miss the reasons behind the number.
Who collected it?
- Primary data is collected first-hand by the researcher, for this study. Experiments, questionnaires, interviews and observations give primary data. It fits the aim exactly, but takes time and money.
- Secondary data was collected by someone else before. Government statistics, a census or an older study are examples. It is cheap and fast, but it may be out of date or not fit the aim.
Mean, median, mode and range
An average (measure of central tendency) gives one typical value.
- Mean: add all values, divide by how many there are. It uses every value, but one very large or small value (an outlier) can pull it away.
- Median: put the values in order and take the middle one. With an even count, take the mean of the two middle values. Outliers do not affect it much.
- Mode: the value that appears most often. There can be no mode or more than one. It is the only average for category data such as "favourite colour".
A measure of dispersion tells how spread out data is. The range = largest value โ smallest value. It is quick to find, but one outlier can make it huge. (Higher levels also use the standard deviation.)
Tables, bar charts and scatter diagrams
A table lists results in rows and columns, often with the mean of each group. A bar chart compares separate groups or conditions: bars do not touch and the x-axis has categories. A histogram shows continuous data: bars touch. A scatter diagram plots two number variables for each person, one on each axis.
- Dots rising left to right: positive correlation (more sleep, higher score).
- Dots falling: negative correlation (more phone time, less sleep).
- No pattern: no correlation.
Remember: correlation does not prove that one thing causes the other.
The normal distribution
Measure enough people's height, IQ score or reaction time and plot how often each value appears. You get a smooth bell-shaped curve. This is the normal distribution.
- It is symmetrical: the left half mirrors the right half.
- The mean, median and mode are all at the centre peak.
- Most values lie near the middle; very few lie at the two ends (the tails).
If the tail is longer on one side the data is skewed: then mean, median and mode separate, and the median is usually the fairer average.
Try it at home
Ask 7 family members or friends how many hours they slept last night. Write the numbers in order. Find the mean, median, mode and range. Then ask how they felt in the morning and write their words: that is your qualitative data. Use the sliders in the 3D (last step) to check your averages.
Key formulas and definitions
- Mean = sum of all values รท number of values
- Median = middle value when data is in order (even count: mean of the two middle values)
- Mode = most frequent value
- Range = largest value โ smallest value
- Normal distribution: symmetrical bell curve, mean = median = mode
Worked examples
1. Classify: (a) number of words recalled, (b) an interview answer about feelings, (c) hospital records used in a study.
(a) Quantitative (a number). (b) Qualitative (words). (c) Secondary data, because the hospital collected it before the study.
2. Recall scores: 6, 9, 7, 6, 12. Find the mean, median, mode and range.
Sum = 40, so mean = 40 รท 5 = 8. In order: 6, 6, 7, 9, 12, so median = 7. Mode = 6. Range = 12 โ 6 = 6.
3. Reaction times (s): 0.4, 0.5, 0.5, 0.6, 0.6, 0.7. Find the median.
There are 6 values, so take the two middle ones: 0.5 and 0.6. Median = (0.5 + 0.6) รท 2 = 0.55 s.
4. Scores: 5, 6, 6, 7, 46. Which average describes the group best and why?
Mean = 70 รท 5 = 14, which is higher than four of the five scores because of the outlier 46. The median (6) is a better typical value.
5. A scatter diagram of hours of revision vs test score shows dots rising to the right. What does it show? Can we say revision causes higher scores?
A positive correlation: more revision goes with higher scores. It does not prove cause; another factor (such as interest in the subject) could affect both.
6. In a large group the mean IQ is 100, and the scores form a normal distribution. What are the median and mode? Where do most scores lie?
In a normal distribution mean = median = mode, so both are 100. Most scores lie close to 100; very few are far above or below.
Common mistakes
- Finding the median without putting the data in order first.
- Thinking quantitative data is always better. Qualitative data explains the reasons behind behaviour.
- Calling data primary just because it is numbers. Primary/secondary is about who collected it, not its form.
- Saying a correlation proves cause. It only shows two variables change together.