What is big data?
Data means facts we can record: numbers, words, pictures, sounds. Most data fits easily on one computer.
Big data is different. It is data that is too large, arrives too fast, or is too mixed to store and study on one normal computer with normal software like a spreadsheet.
We often describe big data with three words that start with V:
- Volume: the amount. Think terabytes (TB) and petabytes (PB). 1 PB = 1000 TB.
- Velocity: the speed at which new data arrives and must be handled, often in real time.
- Variety: the many kinds of data. Some is structured (rows and columns), much is unstructured (photos, video, chat messages).
Some books add more Vs, such as Veracity (can we trust it?) and Value (is it useful?).
Data science is the skill of collecting, cleaning, studying and explaining data to answer questions. Big data is one of the main things data scientists work with.
Why one computer is not enough: processing big data
A relational database keeps data in tables with a fixed design. Big data is often unstructured and keeps changing, so it does not fit well into fixed tables.
Also, one machine has limited memory and speed. So we use distributed processing: the data and the work are spread over many computers (a cluster) that run in parallel, at the same time.
MapReduce in three moves
- Split: cut the data into chunks, one per machine.
- Map: each machine works on its own chunk and produces small results (for example, word → count).
- Reduce: results with the same key are brought together and combined into the final answer.
Example: to count how often each word appears in a billion web pages, each machine counts words in its own pages; then all counts for 'cricket' are added together.
Why functional programming helps
Functions that have no side effects and use immutable data (data that is never changed in place) always give the same output for the same input. So it is safe to run them on many machines at the same time in any order. Higher-order functions like map, filter and reduce match this style exactly.
Modelling big data: facts and graphs
Fact-based model
Instead of changing a record, we store each piece of information as a fact that is never deleted, with a time stamp. Example: 'Asha lives in Pune (2024-03-01)', later 'Asha lives in Delhi (2026-07-15)'. Nothing is overwritten, so we keep the full history and mistakes can be corrected by adding new facts.
Graph schema
A graph stores data as nodes (things such as people, places, products) and edges (links between them, such as 'follows' or 'bought'). Nodes and edges can have properties (for example, age or date). Graphs suit social networks, maps and recommendations because the links matter as much as the things.
Uses, AI and risks of big data
- Science: weather and climate models, telescopes, gene sequencing, particle physics.
- Health: spotting disease outbreaks early from many hospital records.
- Cities: traffic lights and buses planned from sensor data.
- Business: online shops suggest products from what millions bought.
- AI: machine-learning models learn patterns from huge amounts of examples. More good data usually means a better model; biased data means a biased model.
Risks: loss of privacy, data leaks, unfair decisions from biased data, wrong conclusions from 'correlation' that is not 'cause', and the energy used by data centres. Laws such as data-protection acts ask organisations to collect only what they need and keep it safe.
Try it yourself: be a MapReduce team
You need: a newspaper page and 3 or 4 friends or family members.
- Predict: will one person or four people count the word 'the' faster?
- One person counts every 'the' on the whole page. Time it.
- Now cut the page into 4 strips (Split). Each person counts on their own strip (Map).
- Add the four numbers (Reduce). Time it again.
- Compare. Why is the team not exactly 4 times faster? (Cutting, passing and adding also take time.)
In the 3D, use the sliders on the last step to see the same idea.
Key formulas and definitions
- Big data 3 Vs: Volume (amount), Velocity (speed), Variety (kinds); extra: Veracity (trust), Value (use)
- Units: 1 KB = 1000 B, 1 MB = 1000 KB, 1 GB = 1000 MB, 1 TB = 1000 GB, 1 PB = 1000 TB
- Ideal parallel time ≈ time on 1 machine ÷ number of machines (real time is a bit more)
- MapReduce: Split → Map (key, value) → shuffle by key → Reduce (combine)
- Structured data: fixed rows and columns; unstructured data: images, video, free text
- Graph: nodes (things) + edges (relationships) + properties (details)
Worked examples
1. A video app receives 500 hours of new video every minute, plus comments and likes. Which Vs of big data does this show?
Volume: video files are very large and 500 hours a minute is a huge amount. Velocity: it arrives every minute, non-stop. Variety: video, text comments and number counts (likes) are different kinds of data. So all three Vs.
2. One computer needs 60 hours to process a data set. The work is split equally over 12 machines. Ignoring extra costs, how long will it take?
Ideal parallel time = 60 ÷ 12 = 5 hours. In real life it will be a little more, because splitting the data, sending it over the network and combining results also take time.
3. Use MapReduce to count fruit in three baskets: Basket 1 = apple, mango, apple; Basket 2 = mango, mango; Basket 3 = apple, banana.
Map (each basket on its own): B1 → (apple,2), (mango,1); B2 → (mango,2); B3 → (apple,1), (banana,1). Shuffle by key: apple → [2,1], mango → [1,2], banana → [1]. Reduce (add): apple = 3, mango = 3, banana = 1.
4. Draw (in words) a small graph for: Ravi follows Meena, Meena follows Ravi, Meena likes the post 'Rainy day'.
Nodes: Ravi (person), Meena (person), 'Rainy day' (post). Edges: Ravi → Meena labelled 'follows'; Meena → Ravi labelled 'follows'; Meena → 'Rainy day' labelled 'likes'. A property could be the date on the 'likes' edge.
Common mistakes
- Thinking big data only means 'a lot of data'. Speed (velocity) and mix of types (variety) also make data 'big'.
- Believing more machines always means exactly proportionally faster. Splitting, network transfer and combining add extra time.
- Thinking a fact-based model overwrites old values. It adds new time-stamped facts and keeps the old ones.
- Treating a pattern in big data as proof of cause. Two things rising together (correlation) does not mean one causes the other.