The data analysis cycle
- Ask a clear question ("Do students who use phones late sleep less?").
- Collect data: surveys, measurements, sensors, experiments or existing datasets.
- Clean the data.
- Organise / transform: sort, group, make tables, change units.
- Analyse: summaries, patterns, comparisons.
- Visualise and conclude: make a chart, answer the question, say how sure you are.
Often the answer raises a new question, so the cycle starts again. Tools range from paper and a calculator to spreadsheets and programming languages such as Python.
Collecting and cleaning data
Quantitative data are numbers (height 152 cm). Qualitative data are words or categories (favourite sport). Primary data you collect yourself; secondary data someone else collected.
A sample is the group you actually ask. A bigger, randomly chosen sample gives fairer results; asking only your friends adds bias.
Cleaning checklist
- Remove duplicates (the same record twice).
- Fix or remove impossible values (25 hours of sleep in a day).
- Decide what to do with missing values (blanks).
- Make units and spelling the same (cm vs m; "India" vs "india").
Keeping data safe
Personal data needs permission, should be stored securely, and names can be removed (anonymised) before sharing.
Analysing data: averages, spread, outliers
Mean = sum of values ÷ number of values. Median = middle value after sorting (for an even count, the mean of the middle two). Mode = most common value. Range = largest − smallest; it shows spread.
An outlier is a value far from the rest. It can be a mistake or a real rare case. The mean is pulled by outliers; the median is not, so the median is better for skewed data such as incomes.
Exploratory data analysis means looking at data from many sides (tables, charts, summaries) before testing an idea. Grouping (for example by class or city) and combining several datasets can reveal new patterns.
Visualising, concluding and evaluating
Choose the chart for the job: bar chart to compare groups, line graph for change over time, scatter plot for two number variables, pie chart for parts of a whole, histogram for how values are spread.
In a scatter plot, points rising together show a positive correlation; one going up while the other goes down shows a negative correlation. Correlation is not causation: ice-cream sales and drownings both rise in summer because of heat, not because of each other.
Evaluate your conclusion
Was the sample big and fair? Were errors cleaned? Could another analysis of the same data give a different answer? State limits honestly. With big data (millions of rows, e.g. city air-pollution sensors), computers do the same steps faster, but the same care is needed.
Key formulas and definitions
- Mean = sum of values ÷ number of values
- Median = middle value of sorted data (even count: mean of the two middle values)
- Mode = most frequent value
- Range = maximum − minimum
- Correlation ≠ causation
Worked examples
1. Clean this sleep data: 7, 8, 8 (same student twice), 25, blank, 6.
Remove the duplicate 8, the impossible 25 and the blank. Clean data: 7, 8, 6.
2. Find the mean and median of 6, 7, 7, 8, 8, 8, 9, 10.
Sum = 63, count = 8, mean = 63 ÷ 8 = 7.875 ≈ 7.9. Middle two values are 8 and 8, so median = 8.
3. Add an outlier 24 to the data above. What happens to the mean and median?
New sum = 87, count = 9, mean = 87 ÷ 9 ≈ 9.67 (a big jump). Sorted, the 5th value is 8, so median = 8 (no change). The median resists outliers.
Common mistakes
- Skipping cleaning. One impossible value can ruin a mean.
- Forgetting to sort the data before finding the median.
- Saying one thing causes another just because they are correlated.
- Using a pie chart for data that does not add up to one whole, or a line graph for categories.