What is data and what is a dataset?
Data are recorded facts: numbers, words, dates, pictures, clicks. On their own they mean little; when we organise and interpret them they become information, and information we understand and use becomes knowledge.
A dataset is a collection of data, usually a table:
- Row (record, case): one person, object or event.
- Column (variable, field): one property measured for every case.
- Types of variables: categorical (bus, bike, walk) or numerical (age 14, distance 2.5 km).
Data processing means turning raw data into useful information: collect → clean (fix typos, remove duplicates, handle missing values) → sort and filter → calculate (counts, averages, percentages) → visualise (tables, bar charts, line graphs) → interpret. Spreadsheets and simple programs do most of this.
Why a dataset shows a limited picture
No dataset is the whole truth. Always ask what is missing:
- Who is included? A sample is part of the population. If the sample is not chosen fairly (only one school, only people who answer online polls), the result is biased.
- How big is the sample? Small samples give results that jump around by chance. Bigger, random samples are more reliable.
- What was measured, and how? The wording of a question, the sensor used, or who collected the data can change the answers.
- When? Data from five years ago may be out of date.
- What is left out? Variables that were not recorded can hide the real cause.
Correlation is not causation: ice-cream sales and swimming accidents both rise in summer, but ice cream does not cause accidents; hot weather drives both.
Charts can also mislead: a bar chart whose axis starts at 90 instead of 0 makes a small difference look huge.
Doing research with a dataset
- Ask a clear question that data can answer: "What share of students in our class walk to school?"
- Find or collect data: a survey, a measurement, or an existing (open) dataset. Note the source.
- Clean it: remove impossible values (age 140), fix spelling ("bus", "Bus", "BUS"), decide what to do with blanks.
- Process it: sort, filter (only Year 9), group and count, work out percentages and averages.
- Visualise: choose the right chart: bar chart for categories, line graph for change over time, pie chart for parts of a whole, scatter graph for two numerical variables.
- Conclude carefully: say what the data shows, how sure you are, and what its limits are.
Worked example: 20 students answered: walk 8, bike 6, bus 4, car 2. Share who walk = 8 ÷ 20 × 100% = 40%. Conclusion: "In this sample of 20, 40% walk. We asked only one class, so the whole school may differ."
Open data
Open data is data that anyone can access, use and share for free, usually under an open licence that asks only for credit to the source. Governments, cities, universities and scientists publish open data on transport, weather, budgets, health, schools and the environment.
Why it matters: it lets citizens check how public money is spent, helps developers build useful apps (bus-arrival apps, pollution maps), and lets students do real research.
Protecting people: before publishing, personal details are removed or grouped (anonymisation): names deleted, exact addresses replaced by area, ages put into ranges. But combining several datasets can sometimes re-identify people, so publishers must be careful.
Good open data is findable, machine-readable (CSV, not a photo of a table), documented (what each column means) and up to date.
How organisations use data
Shops, apps, banks, hospitals and governments collect data, often personal data (anything that can identify you: name, phone number, location, face, purchase history).
- Businesses analyse purchases to stock shelves, set prices and show targeted adverts based on your profile.
- Streaming and social apps use your clicks and viewing time to recommend content and keep you watching.
- Governments use census and traffic data to plan schools, hospitals and roads.
- Hospitals and scientists use health data to find better treatments.
Benefits: better services, less waste, useful recommendations. Risks: loss of privacy, data leaks, unfair decisions from biased data, manipulation by personalised content.
Your rights: many countries have data-protection laws (for example the EU's GDPR and India's Digital Personal Data Protection Act, 2023). Common ideas: organisations must have a lawful reason or your consent, collect only what they need, keep it safe, and let you see or delete your data.
Try it: a mini survey at home
Ask 10 family members or friends one question, such as "How did you travel today?" Write a table with one row per person and columns for age group and answer. Count each answer, draw a bar chart and work out the percentages. Then write one sentence on what your data cannot tell you (who you did not ask). Compare with the free-play step: is a sample of 10 big enough?
Key formulas and definitions
- Row = one case (record); Column = one variable (field)
- Data → (process) → Information → (understand) → Knowledge
- Percentage = part ÷ total × 100%
- Mean (average) = sum of values ÷ number of values
- Sample = part of the population that is actually measured
- Bias = a systematic error that pushes results one way
- Open data = free to access, use and share (with an open licence)
- Personal data = any data that can identify a person
Worked examples
1. A dataset has 250 rows and 8 columns. How many cases and how many variables?
250 cases (one per row) and 8 variables (one per column). Total cells = 250 × 8 = 2000.
2. Of 40 students surveyed, 14 cycle to school. What percentage cycle?
14 ÷ 40 × 100% = 35%.
3. An online poll on a gaming website finds that 85% of people play games every day. Can we say 85% of the country does?
No. The sample is biased: only visitors to a gaming site answered, and they chose to answer. The result shows only that group.
4. A city publishes a table of bus stops with name, location and number of passengers per day. Is this open data, and does it contain personal data?
Yes, it is open data if anyone can download and reuse it. It has no personal data, because it counts passengers without identifying anyone.
Common mistakes
- Treating a small or one-sided sample as if it speaks for everyone.
- Thinking that a link between two variables proves one causes the other.
- Skipping cleaning, so typos split one category into several ("Bus" and "bus").
- Believing anonymised data can never be traced back. Combining datasets can re-identify people.