Population, sample and census
The population is the whole group we want to study: all students in a school, all bulbs from a factory, all fish in a lake.
A census checks every member. It is exact but slow and costly, and sometimes impossible (you cannot test every match-stick by burning it!).
A sample is a smaller part of the population that we actually check. We use the sample to estimate facts about the population. This step from sample to population is called statistical inference.
Simple random sampling
In a simple random sample, every member of the population has the same chance of being chosen, and every group of the chosen size is equally likely.
How to do it: number every member 1 to N. Then use a lottery (slips in a box), a random-number table, or a random-number function on a calculator, spreadsheet (RAND) or in Python (random.sample). Skip repeats.
Stratified and systematic sampling
Stratified sampling: split the population into groups (strata) that differ, such as classes, age bands or regions. Take a random sample from each group, in proportion to its size.
Example: a school has 600 boys and 400 girls. For a sample of 50: boys = 600/1000 × 50 = 30, girls = 20.
Systematic sampling: from a list of N, choose every k-th member, where k = N ÷ n. Start at a random place between 1 and k. Easy, but it can go wrong if the list has a repeating pattern.
Bias: when a sample misleads
A sample is biased if some members are more likely to be chosen, so the sample does not look like the population.
- Convenience sample: asking only friends or the people near you.
- Self-selection: only people who care strongly reply to an online poll.
- Non-response: many chosen people do not answer.
- Leading questions: "Don't you agree that homework is too long?"
A bigger biased sample is still wrong. Fix bias by choosing at random, not by adding more people.
Sampling variation and the law of large numbers
Two random samples from the same population usually give slightly different answers. This is sampling variation (or fluctuation). It is not a mistake; it is chance.
Law of large numbers: as the sample gets bigger, the sample proportion gets closer to the true proportion p.
A handy rule: for a sample of size n, about 95% of samples give a proportion within p ± 1/√n. With n = 100 that is ±0.1; with n = 400 it is ±0.05. To halve the wobble you need four times the sample.
Simulation: repeating random trials
A simulation uses random numbers to copy a chance process many times. Example: to model a coin, generate a random number between 0 and 1 and call it heads if it is below 0.5.
In a spreadsheet: =IF(RAND()<0.3,1,0) gives 1 with probability 0.3. Copy it into 100 cells and average them: you get one sample proportion. Repeat to see how the proportions spread around 0.3.
The 3D free play is exactly such a simulation.
Estimating from a sample
If a sample proportion is p̂, estimate the number in the population as p̂ × N.
Capture–recapture (for animals): tag M animals and release them. Later catch n, of which m are tagged. Then m/n ≈ M/N, so N ≈ M × n ÷ m.
Results of a survey are often shown with a pie chart or a frequency histogram, and always with the sample size, so readers can judge how reliable they are.
Try it: a sweet-jar survey
Put 30 red and 70 other small items (beads, buttons) in a bag. Without looking, take 10, count the red, put them back, shake. Do this 5 times and write each %. Then take 40 at a time, 5 times. Which set of answers stays closer to 30%? Now try the same in the 3D free play with sizes 10 and 200.
Key formulas and definitions
- Sample proportion p̂ = (number with the feature) ÷ (sample size n)
- Estimate in population ≈ p̂ × N
- Stratified share for a group = (group size ÷ N) × n
- Systematic step k = N ÷ n; about 95% of p̂ lie in p ± 1/√n
- Capture–recapture: N ≈ M × n ÷ m
Worked examples
1. A factory wants the average life of its bulbs and tests 50 of them. Name the population and the sample. Why not a census?
Population: all bulbs the factory makes. Sample: the 50 tested bulbs. A census would burn out every bulb, leaving nothing to sell.
2. In a random sample of 40 students, 12 walk to school. The school has 900 students. Estimate how many walk.
p̂ = 12 ÷ 40 = 0.3. Estimate = 0.3 × 900 = 270 students.
3. A school has 600 boys and 400 girls. Choose a stratified sample of 50.
Boys: 600/1000 × 50 = 30. Girls: 400/1000 × 50 = 20. Pick each group at random.
4. From a list of 500 members, take a systematic sample of 25.
k = 500 ÷ 25 = 20. Pick a random start from 1 to 20, say 7. Then take 7, 27, 47, … up to 487.
5. 50 fish are tagged and released. Later 40 fish are caught; 5 are tagged. Estimate the fish in the pond.
N ≈ 50 × 40 ÷ 5 = 400 fish.
6. The true share is p = 0.3 and you take samples of 100. Between which values do about 95% of sample proportions fall?
1/√100 = 0.1, so between 0.3 − 0.1 = 0.2 and 0.3 + 0.1 = 0.4.
7. A website asks visitors to vote on whether people read enough books. 5000 vote. Is this a good sample of the country?
No. It is self-selected: only website visitors who chose to vote are included. It is biased, so a large size does not help.
Common mistakes
- Thinking a bigger sample fixes bias. It only reduces random wobble, not bias.
- Calling any unplanned choice "random". Random means every member has an equal, known chance.
- In stratified sampling, taking the same number from each group even when the groups are different sizes.
- Expecting every random sample to give the exact true value. Some variation is normal.