📘 CodingMarble Learn

Sampling: Learning About a Population from a Sample

A population is the whole group we want to know about; a sample is a smaller part we actually check. A good sample is chosen at random so that it represents the population. Simple random sampling gives everyone an equal chance; stratified sampling takes the right share from each group; systematic sampling takes every k-th item. Different samples give slightly different answers (sampling variation), but bigger samples wobble less (law of large numbers). A biased sample gives a wrong answer however big it is.

🎬 Step-by-step story

  1. Here is the whole population: 400 students in 4 classes. Orange ones walk to school. Asking all 400 takes too long.
  2. We pick 20 students at random. Each student has the same chance. The share of orange in the sample is our estimate.
  3. We take a second random sample of 20. The estimate changes a little. This natural wobble is called sampling variation.
  4. Now we take a sample of 100. Its estimate lands closer to the true 30%. Bigger random samples wobble less.
  5. This time we pick 20 only from class A, where many students live nearby. The estimate is far too high. That sample is biased.
  6. Free play: set the sample size, take many random, stratified or biased samples, and see where the dots land on the 0–100% line.

Tip: drag the 3D scene to turn it. Use two fingers to zoom.

🤔 Common doubts, cleared

Why not just ask everyone?

A census is slow, costly and sometimes impossible. Step 1 shows how big even a small school population is.

How can 20 people tell us about 400?

If chosen at random, the sample is a small copy of the population. Step 2 shows the estimate from 20.

My two samples gave different answers. Did I do it wrong?

No. That is sampling variation. Step 3 takes two samples and gets two answers.

Why are bigger samples better?

They wobble less around the truth (law of large numbers). Compare sample 100 in step 4 with sample 20.

Will a huge sample fix a bad choosing method?

No. Step 5 shows a biased sample; in free play pick 'Class A only' with size 100 and it is still too high.

Why does stratified sampling help?

It makes sure every group is fairly included. In free play the purple stratified dots cluster tightly around 30%.

Population, sample and census

The population is the whole group we want to study: all students in a school, all bulbs from a factory, all fish in a lake.

A census checks every member. It is exact but slow and costly, and sometimes impossible (you cannot test every match-stick by burning it!).

A sample is a smaller part of the population that we actually check. We use the sample to estimate facts about the population. This step from sample to population is called statistical inference.

Simple random sampling

In a simple random sample, every member of the population has the same chance of being chosen, and every group of the chosen size is equally likely.

How to do it: number every member 1 to N. Then use a lottery (slips in a box), a random-number table, or a random-number function on a calculator, spreadsheet (RAND) or in Python (random.sample). Skip repeats.

Stratified and systematic sampling

Stratified sampling: split the population into groups (strata) that differ, such as classes, age bands or regions. Take a random sample from each group, in proportion to its size.

Example: a school has 600 boys and 400 girls. For a sample of 50: boys = 600/1000 × 50 = 30, girls = 20.

Systematic sampling: from a list of N, choose every k-th member, where k = N ÷ n. Start at a random place between 1 and k. Easy, but it can go wrong if the list has a repeating pattern.

Bias: when a sample misleads

A sample is biased if some members are more likely to be chosen, so the sample does not look like the population.

A bigger biased sample is still wrong. Fix bias by choosing at random, not by adding more people.

Sampling variation and the law of large numbers

Two random samples from the same population usually give slightly different answers. This is sampling variation (or fluctuation). It is not a mistake; it is chance.

Law of large numbers: as the sample gets bigger, the sample proportion gets closer to the true proportion p.

A handy rule: for a sample of size n, about 95% of samples give a proportion within p ± 1/√n. With n = 100 that is ±0.1; with n = 400 it is ±0.05. To halve the wobble you need four times the sample.

Simulation: repeating random trials

A simulation uses random numbers to copy a chance process many times. Example: to model a coin, generate a random number between 0 and 1 and call it heads if it is below 0.5.

In a spreadsheet: =IF(RAND()<0.3,1,0) gives 1 with probability 0.3. Copy it into 100 cells and average them: you get one sample proportion. Repeat to see how the proportions spread around 0.3.

The 3D free play is exactly such a simulation.

Estimating from a sample

If a sample proportion is p̂, estimate the number in the population as p̂ × N.

Capture–recapture (for animals): tag M animals and release them. Later catch n, of which m are tagged. Then m/n ≈ M/N, so N ≈ M × n ÷ m.

Results of a survey are often shown with a pie chart or a frequency histogram, and always with the sample size, so readers can judge how reliable they are.

Try it: a sweet-jar survey

Put 30 red and 70 other small items (beads, buttons) in a bag. Without looking, take 10, count the red, put them back, shake. Do this 5 times and write each %. Then take 40 at a time, 5 times. Which set of answers stays closer to 30%? Now try the same in the 3D free play with sizes 10 and 200.

Key formulas and definitions

Worked examples

1. A factory wants the average life of its bulbs and tests 50 of them. Name the population and the sample. Why not a census?

Population: all bulbs the factory makes. Sample: the 50 tested bulbs. A census would burn out every bulb, leaving nothing to sell.

2. In a random sample of 40 students, 12 walk to school. The school has 900 students. Estimate how many walk.

p̂ = 12 ÷ 40 = 0.3. Estimate = 0.3 × 900 = 270 students.

3. A school has 600 boys and 400 girls. Choose a stratified sample of 50.

Boys: 600/1000 × 50 = 30. Girls: 400/1000 × 50 = 20. Pick each group at random.

4. From a list of 500 members, take a systematic sample of 25.

k = 500 ÷ 25 = 20. Pick a random start from 1 to 20, say 7. Then take 7, 27, 47, … up to 487.

5. 50 fish are tagged and released. Later 40 fish are caught; 5 are tagged. Estimate the fish in the pond.

N ≈ 50 × 40 ÷ 5 = 400 fish.

6. The true share is p = 0.3 and you take samples of 100. Between which values do about 95% of sample proportions fall?

1/√100 = 0.1, so between 0.3 − 0.1 = 0.2 and 0.3 + 0.1 = 0.4.

7. A website asks visitors to vote on whether people read enough books. 5000 vote. Is this a good sample of the country?

No. It is self-selected: only website visitors who chose to vote are included. It is biased, so a large size does not help.

Common mistakes

Practice quiz

1. A census means:
2. In simple random sampling, every member has:
3. Splitting into groups and sampling each in proportion is:
4. As a random sample gets bigger, its proportion:
5. Asking only your friends is an example of:

Practice: answer these yourself

Type or choose your answer, then press Check. Use a hint if you are stuck; the full solution appears after you answer.

Frequently asked questions

What is the difference between a population and a sample?

The population is the whole group you want to know about. A sample is the smaller part you actually measure.

What are the main types of sampling?

Simple random, stratified, systematic and cluster sampling are random methods. Convenience and self-selected samples are non-random and often biased.

Why is random sampling important?

It gives every member an equal chance, so the sample represents the population and the size of the error can be estimated.

Where this is taught

Spain2º ESOStochastic sense
Spain3º ESOStochastic sense
Spain4º ESOStochastic sense
Spain4º ESOStochastic sense
Spain1º BachilleratoStochastic sense
CBSE (India)Class 12Inferential Statistics
England (GCSE, A level)Year 101. The collection of data
England (GCSE, A level)Year 112. Processing, representing and analysing data (part 2)
England (GCSE, A level)Year 12K-L Statistical sampling and data
England (GCSE, A level)Year 133.7 Research methods
USA (Common Core, NGSS, AP)Grade 11Inferences and conclusions from data
USA (Common Core, NGSS, AP)Grade 11Inferences and conclusions from data
USA (Common Core, NGSS, AP)Grade 12Exploring One-Variable Data and Collecting Data
Japan中学3年Using data
Japan高校2年Statistical inference
South Korea고등학교 2학년Preparing and analysing data
South Korea고등학교 2학년Statistics and statistical problems
South Korea고등학교 2학년Statistics
South Korea고등학교 3학년Statistics
FranceSecondeStatistics and probability
FrancePremièreProbability and statistics
China九年级(初三)Probability and statistics (standard, placement varies)
China高一Ch.9 Statistics
China高三Elective C (humanities)

Learn first

Learn next

Related lessons

All Maths lessons