📘 CodingMarble Learn

Big Data

Big data is data that is too big, too fast or too mixed to store and process on one ordinary computer with normal tools. We describe it with Volume (how much), Velocity (how fast it arrives) and Variety (how many kinds). To handle it, the work is split across many machines that run at the same time (distributed processing, for example MapReduce). Big data is often stored as simple facts or as a graph of nodes and links, and it is used for weather forecasts, maps, health, shopping and training AI. It also raises questions about privacy and fairness.

🎬 Step-by-step story

  1. Here is one computer and a small amount of data: 8 blocks. One machine can store and sort this easily.
  2. Volume: the data keeps growing until the pile is far bigger than one machine can hold.
  3. Velocity: new data does not wait. It flows in every second, like a stream from phones and sensors.
  4. Variety: the data comes as numbers, photos, videos and text. It does not fit one neat table.
  5. Split and combine: the job is cut into parts and sent to 4 machines at once (Map). Their answers are joined (Reduce). The job finishes about 4 times faster.
  6. Try it: change the amount of data and the number of machines. See how the time changes.

Tip: drag the 3D scene to turn it. Use two fingers to zoom.

🤔 Common doubts, cleared

How big does data have to be to count as 'big data'?

There is no fixed number. Data is 'big' when one normal computer and normal tools cannot store or process it in useful time. Step 2 shows the pile outgrowing one machine.

Why can't we just buy one bigger computer?

One machine has a limit and gets very expensive. Adding many ordinary machines is cheaper and can grow without limit, as step 5 shows.

Is fast data really a problem if it is small?

Yes. If data arrives faster than you can process it, it piles up. Step 3 shows a non-stop stream that must be handled in real time.

Why don't 4 machines make it exactly 4 times faster?

Splitting data, sending it over the network and joining results take extra time. Try changing the machines slider on the last step.

Why does variety make data hard?

A photo, a number and a sentence cannot sit in the same neat column. Step 4 shows the different shapes.

What is big data?

Data means facts we can record: numbers, words, pictures, sounds. Most data fits easily on one computer.

Big data is different. It is data that is too large, arrives too fast, or is too mixed to store and study on one normal computer with normal software like a spreadsheet.

We often describe big data with three words that start with V:

Some books add more Vs, such as Veracity (can we trust it?) and Value (is it useful?).

Data science is the skill of collecting, cleaning, studying and explaining data to answer questions. Big data is one of the main things data scientists work with.

Why one computer is not enough: processing big data

A relational database keeps data in tables with a fixed design. Big data is often unstructured and keeps changing, so it does not fit well into fixed tables.

Also, one machine has limited memory and speed. So we use distributed processing: the data and the work are spread over many computers (a cluster) that run in parallel, at the same time.

MapReduce in three moves

  1. Split: cut the data into chunks, one per machine.
  2. Map: each machine works on its own chunk and produces small results (for example, word → count).
  3. Reduce: results with the same key are brought together and combined into the final answer.

Example: to count how often each word appears in a billion web pages, each machine counts words in its own pages; then all counts for 'cricket' are added together.

Why functional programming helps

Functions that have no side effects and use immutable data (data that is never changed in place) always give the same output for the same input. So it is safe to run them on many machines at the same time in any order. Higher-order functions like map, filter and reduce match this style exactly.

Modelling big data: facts and graphs

Fact-based model

Instead of changing a record, we store each piece of information as a fact that is never deleted, with a time stamp. Example: 'Asha lives in Pune (2024-03-01)', later 'Asha lives in Delhi (2026-07-15)'. Nothing is overwritten, so we keep the full history and mistakes can be corrected by adding new facts.

Graph schema

A graph stores data as nodes (things such as people, places, products) and edges (links between them, such as 'follows' or 'bought'). Nodes and edges can have properties (for example, age or date). Graphs suit social networks, maps and recommendations because the links matter as much as the things.

Uses, AI and risks of big data

Risks: loss of privacy, data leaks, unfair decisions from biased data, wrong conclusions from 'correlation' that is not 'cause', and the energy used by data centres. Laws such as data-protection acts ask organisations to collect only what they need and keep it safe.

Try it yourself: be a MapReduce team

You need: a newspaper page and 3 or 4 friends or family members.

  1. Predict: will one person or four people count the word 'the' faster?
  2. One person counts every 'the' on the whole page. Time it.
  3. Now cut the page into 4 strips (Split). Each person counts on their own strip (Map).
  4. Add the four numbers (Reduce). Time it again.
  5. Compare. Why is the team not exactly 4 times faster? (Cutting, passing and adding also take time.)

In the 3D, use the sliders on the last step to see the same idea.

Key formulas and definitions

Worked examples

1. A video app receives 500 hours of new video every minute, plus comments and likes. Which Vs of big data does this show?

Volume: video files are very large and 500 hours a minute is a huge amount. Velocity: it arrives every minute, non-stop. Variety: video, text comments and number counts (likes) are different kinds of data. So all three Vs.

2. One computer needs 60 hours to process a data set. The work is split equally over 12 machines. Ignoring extra costs, how long will it take?

Ideal parallel time = 60 ÷ 12 = 5 hours. In real life it will be a little more, because splitting the data, sending it over the network and combining results also take time.

3. Use MapReduce to count fruit in three baskets: Basket 1 = apple, mango, apple; Basket 2 = mango, mango; Basket 3 = apple, banana.

Map (each basket on its own): B1 → (apple,2), (mango,1); B2 → (mango,2); B3 → (apple,1), (banana,1). Shuffle by key: apple → [2,1], mango → [1,2], banana → [1]. Reduce (add): apple = 3, mango = 3, banana = 1.

4. Draw (in words) a small graph for: Ravi follows Meena, Meena follows Ravi, Meena likes the post 'Rainy day'.

Nodes: Ravi (person), Meena (person), 'Rainy day' (post). Edges: Ravi → Meena labelled 'follows'; Meena → Ravi labelled 'follows'; Meena → 'Rainy day' labelled 'likes'. A property could be the date on the 'likes' edge.

Common mistakes

Practice quiz

1. Which of these is NOT one of the original 3 Vs of big data?
2. In MapReduce, the step that combines results with the same key is:
3. Photos, videos and chat messages are examples of:
4. Why is functional programming useful for distributed processing?
5. In a graph schema, a relationship such as 'follows' is stored as:

Practice: answer these yourself

Type or choose your answer, then press Check. Use a hint if you are stuck; the full solution appears after you answer.

Frequently asked questions

What is big data in simple words?

Data that is too big, too fast or too mixed for one ordinary computer, so many computers have to work on it together.

What are the 3 Vs of big data?

Volume (how much), Velocity (how fast) and Variety (how many kinds). Some add Veracity and Value.

What is MapReduce?

A way to process big data: split it, let many machines work on parts (Map), then combine their results (Reduce).

Where this is taught

CBSE (India)Class 12Big Data and Data Analytics
England (GCSE, A level)Year 134.11 Big Data
South Korea고등학교 1학년Science and the future
South Korea고등학교 2학년AI and big data
South Korea고등학교 2학년Data
China高一Comp.1 Ch.1 Data and big data

Learn first

Learn next

Related lessons

All Computer Science lessons