📘 CodingMarble Learn

Data Literacy: Reading, Using and Questioning Data

Data literacy means being able to read data, work with it and ask good questions about it. A dataset is a table: each row is a case (a person, a day, a shop) and each column is a variable (age, travel mode, price). Every dataset shows only a limited picture: it covers some cases, some variables, some time and was collected in a certain way, so it can be biased or out of date. Researching with a dataset follows steps: ask a question, find or collect data, clean it, sort, filter and count, make a chart, draw a careful conclusion. Open data is data anyone may use for free, usually anonymised first. Organisations collect data about us (purchases, location, clicks) to build profiles, target adverts and plan services, which brings benefits but also privacy risks and rules such as consent.

🎬 Step-by-step story

  1. A dataset is a table. Each row is one case, like one student. Each column is one fact, like age or how they travel.
  2. The town has 100 students. Only 20 from one school were asked. The grey ones are not in the data. So the data shows only part of the picture.
  3. Ask a question: how do students get to school? Count each answer. The counts become a bar chart: walk 40, bike 30, bus 20, car 10.
  4. Open data is free for everyone. Before sharing, the name column is removed. Now nobody can be picked out.
  5. Every small purchase is a record. Many records join into one profile. The profile decides which advert you see.
  6. Your turn. Change the sample size. With few students the guess jumps around. With more, it gets close to the true 40%.

Tip: drag the 3D scene to turn it. Use two fingers to zoom.

🤔 Common doubts, cleared

What is the difference between data and information?

Data are raw facts; information is data organised to answer a question, like the counts in a bar chart. Step 2 turns rows into a chart.

Why can't one school speak for the whole town?

Students at one school live in one area and may travel differently. Step 1 shows 80 students missing from the data.

Is open data safe if it is about people?

Only after names and other identifying details are removed. Step 3 shows the name column fading out.

How does an app know which advert to show me?

It joins many small records about you into a profile. Step 4 shows purchases flying into one profile block.

How big should a sample be?

Big enough that the answer stops jumping around. In free play, move from 10 to 100 and watch the gap shrink.

What is data and what is a dataset?

Data are recorded facts: numbers, words, dates, pictures, clicks. On their own they mean little; when we organise and interpret them they become information, and information we understand and use becomes knowledge.

A dataset is a collection of data, usually a table:

Data processing means turning raw data into useful information: collect → clean (fix typos, remove duplicates, handle missing values) → sort and filter → calculate (counts, averages, percentages) → visualise (tables, bar charts, line graphs) → interpret. Spreadsheets and simple programs do most of this.

Why a dataset shows a limited picture

No dataset is the whole truth. Always ask what is missing:

Correlation is not causation: ice-cream sales and swimming accidents both rise in summer, but ice cream does not cause accidents; hot weather drives both.

Charts can also mislead: a bar chart whose axis starts at 90 instead of 0 makes a small difference look huge.

Doing research with a dataset

  1. Ask a clear question that data can answer: "What share of students in our class walk to school?"
  2. Find or collect data: a survey, a measurement, or an existing (open) dataset. Note the source.
  3. Clean it: remove impossible values (age 140), fix spelling ("bus", "Bus", "BUS"), decide what to do with blanks.
  4. Process it: sort, filter (only Year 9), group and count, work out percentages and averages.
  5. Visualise: choose the right chart: bar chart for categories, line graph for change over time, pie chart for parts of a whole, scatter graph for two numerical variables.
  6. Conclude carefully: say what the data shows, how sure you are, and what its limits are.

Worked example: 20 students answered: walk 8, bike 6, bus 4, car 2. Share who walk = 8 ÷ 20 × 100% = 40%. Conclusion: "In this sample of 20, 40% walk. We asked only one class, so the whole school may differ."

Open data

Open data is data that anyone can access, use and share for free, usually under an open licence that asks only for credit to the source. Governments, cities, universities and scientists publish open data on transport, weather, budgets, health, schools and the environment.

Why it matters: it lets citizens check how public money is spent, helps developers build useful apps (bus-arrival apps, pollution maps), and lets students do real research.

Protecting people: before publishing, personal details are removed or grouped (anonymisation): names deleted, exact addresses replaced by area, ages put into ranges. But combining several datasets can sometimes re-identify people, so publishers must be careful.

Good open data is findable, machine-readable (CSV, not a photo of a table), documented (what each column means) and up to date.

How organisations use data

Shops, apps, banks, hospitals and governments collect data, often personal data (anything that can identify you: name, phone number, location, face, purchase history).

Benefits: better services, less waste, useful recommendations. Risks: loss of privacy, data leaks, unfair decisions from biased data, manipulation by personalised content.

Your rights: many countries have data-protection laws (for example the EU's GDPR and India's Digital Personal Data Protection Act, 2023). Common ideas: organisations must have a lawful reason or your consent, collect only what they need, keep it safe, and let you see or delete your data.

Try it: a mini survey at home

Ask 10 family members or friends one question, such as "How did you travel today?" Write a table with one row per person and columns for age group and answer. Count each answer, draw a bar chart and work out the percentages. Then write one sentence on what your data cannot tell you (who you did not ask). Compare with the free-play step: is a sample of 10 big enough?

Key formulas and definitions

Worked examples

1. A dataset has 250 rows and 8 columns. How many cases and how many variables?

250 cases (one per row) and 8 variables (one per column). Total cells = 250 × 8 = 2000.

2. Of 40 students surveyed, 14 cycle to school. What percentage cycle?

14 ÷ 40 × 100% = 35%.

3. An online poll on a gaming website finds that 85% of people play games every day. Can we say 85% of the country does?

No. The sample is biased: only visitors to a gaming site answered, and they chose to answer. The result shows only that group.

4. A city publishes a table of bus stops with name, location and number of passengers per day. Is this open data, and does it contain personal data?

Yes, it is open data if anyone can download and reuse it. It has no personal data, because it counts passengers without identifying anyone.

Common mistakes

Practice quiz

1. In a dataset table, each row usually stands for…
2. Asking only students from one school about the whole town gives a…
3. Which chart best shows how temperature changes over a week?
4. Removing names before publishing data is called…
5. A shop uses your past purchases to choose adverts for you. This is…

Practice: answer these yourself

Type or choose your answer, then press Check. Use a hint if you are stuck; the full solution appears after you answer.

Frequently asked questions

What is data literacy?

The ability to read, work with, analyse and question data, and to explain what it does and does not show.

What is open data?

Data that anyone can freely access, use and share, usually published by governments or researchers under an open licence.

Why does a dataset show only a limited picture?

Because it covers only some people, variables and times, and the way it was collected can add bias or errors.

Where this is taught

NetherlandsVWO 3 (onderbouw)Data, AI and society
CBSE (India)Class 9Part B: Data Literacy
CBSE (India)Class 10Part B: Identifying Patterns
CBSE (India)Class 10Part B: Data Merging
CBSE (India)Class 11Data Literacy

Learn first

Learn next

Related lessons

All Computer Science lessons