📘 CodingMarble Learn

Statistical Inference

Statistical inference means using a sample to say something about a whole population. A number that describes the population (like the true proportion p or the mean μ) is a parameter. A number worked out from a sample (like p̂ or x̄) is a statistic, and we use it as a point estimate. Different random samples give different answers: this is sampling variability. If we took many samples, their statistics would form the sampling distribution, centred on the true value, with spread called the standard error: SE = √(p(1−p)/n) for a proportion and σ/√n for a mean. Bigger samples give smaller spread. A 95% confidence interval is estimate ± 1.96 × SE; about 95 of every 100 such intervals catch the true value. Simulation helps us check whether a claimed model fits the data. Good inference needs random sampling, and an association in data does not prove cause.

🎬 Step-by-step story

  1. A bag holds 400 beads. We cannot see what share p is red. We pull out 50 beads (a sample) and count the red ones. That share is p̂, our estimate.
  2. We put them back and pull out another 50. A different answer! Every sample gives a slightly different p̂. This is called sampling variability.
  3. Now we take 150 samples and drop each p̂ as a dot. The dots pile into a bell shape. Its middle is near the true p. This pile is the sampling distribution.
  4. Make each sample bigger: n = 200. The pile becomes thinner. Bigger samples wobble less, so we can trust them more.
  5. From each sample we draw a band: p̂ ± 1.96 × SE. This is a 95% confidence interval. About 19 out of 20 bands cross the true p (green). A few miss (red).
  6. Your turn: change the sample size and the true p. Press "Take a sample" many times. Count how many bands catch the true p.

Tip: drag the 3D scene to turn it. Use two fingers to zoom.

🤔 Common doubts, cleared

Why does every sample give a different answer if the bag stays the same?

Each sample picks different beads by chance. Step 1 shows two samples from the same bag giving different p̂.

If samples are random, how can we trust them at all?

Their errors follow a pattern. Step 2 shows many p̂ values piling up around the true p.

Why does a bigger sample help?

The pile of p̂ values gets thinner (SE = √(p(1−p)/n) falls). Step 3 shows n = 200 vs n = 50.

Does a 95% interval always contain the true value?

No. About 1 in 20 misses. Step 4 shows green bands that catch p and red ones that miss.

What is the real difference between p and p̂?

p is the hidden share in the whole bag (fixed). p̂ is the share in your handful (changes). Step 0 shows both.

How do I check if a claim fits my data?

Simulate many samples from the claimed model and see how often a result like yours appears. Use the sliders in step 5.

Population, sample, parameter and statistic

The population is the whole group we want to know about: all voters in a country, all bulbs made today. A sample is the smaller group we actually look at.

Memory trick: Population ↔ Parameter, Sample ↔ Statistic (same first letters).

Why take a sample?

Checking everyone (a census) is often too slow, too costly or impossible (you cannot burn every match to test it). A good sample is much cheaper.

Good samples are random

In a simple random sample, every member has the same chance to be picked. Other methods: stratified (pick randomly inside groups like age bands), systematic (every 10th person), cluster (pick whole classes or villages). Bias appears when some people are more likely to be chosen, for example asking only your friends, or an online poll where only keen people answer.

Sampling variability and the sampling distribution

Two random samples from the same population almost never give the same statistic. This natural wobble is sampling variability. It is not a mistake; it is chance.

If we imagine taking a great many samples of the same size n and plotting every p̂, we get the sampling distribution. For fairly large samples it is:

Because n sits under a square root, making the sample 4 times bigger makes the spread only half as big.

Checking a model by simulation

Suppose a company claims 40% of its customers are students, and your random sample of 50 has only 10 students (20%). Is the claim believable? Use a computer (or the 3D above) to take many fake samples of 50 from a population with p = 0.4. If 20% or less almost never happens, the data are not consistent with the claim. If it happens often, the claim still fits.

Estimating: point estimates and confidence intervals

A point estimate is one number: "p̂ = 0.42". It is our best single guess, but it will almost never be exactly right.

An interval estimate gives a range and a level of trust:

The part after ± is the margin of error. Common values of z*: 1.645 for 90%, 1.96 for 95%, 2.576 for 99%.

What does "95% confident" mean?

If we repeated the sampling many times and built an interval each time, about 95% of the intervals would contain the true value. It does not mean "there is a 95% chance p is in this one interval" in the everyday sense: p is fixed, the interval is what changes (see the green and red bands in step 4).

Making the interval narrower

Sample size for a wanted margin E (proportion, worst case p = 0.5): n ≈ (z*/(2E))². For E = 0.03 at 95%: n ≈ (1.96/0.06)² ≈ 1068.

Comparing groups and judging conclusions

Inference also lets us compare two groups, for example the share of students in two schools who walk to school.

Association is not causation

Ice-cream sales and drowning both rise in summer. They are associated, but ice-cream does not cause drowning; hot weather drives both (a confounding variable). Only a randomised experiment, where chance decides who gets the treatment, can show cause.

Checklist for judging a claim

  1. Was the sample random? Who was left out?
  2. How big was n? What is the margin of error?
  3. Survey, experiment or observational study?
  4. Is the conclusion about the right population?
  5. Is the difference bigger than chance alone would give?

Try it at home

Put 40 red and 60 other-coloured items (buttons, beads, paper slips) into a bag. Shake. Without looking, take 20, write down the share of red, and put them back. Repeat 15 times and make a dot plot. Where is the middle? Now repeat with samples of 5. Which dot plot is wider? Then try the sliders in the 3D above and compare.

Key formulas and definitions

Worked examples

1. In a random sample of 80 students, 28 cycle to school. Find the point estimate of the proportion who cycle.

p̂ = 28 ÷ 80 = 0.35. Our best single estimate is 35%.

2. The true proportion is p = 0.4. Find the standard error of p̂ for samples of n = 100 and n = 400.

n = 100: SE = √(0.4 × 0.6 / 100) = √0.0024 ≈ 0.049. n = 400: SE = √(0.24/400) = √0.0006 ≈ 0.0245. Four times the sample halves the SE.

3. A survey of 500 adults finds 210 use a fitness app. Build a 95% confidence interval for the true proportion.

p̂ = 210/500 = 0.42. SE = √(0.42 × 0.58 / 500) = √0.000487 ≈ 0.0221. Margin = 1.96 × 0.0221 ≈ 0.043. Interval: 0.42 ± 0.043 = (0.377, 0.463). We are 95% confident that between about 37.7% and 46.3% of adults use such an app.

4. Packets of rice have σ = 8 g. A sample of 64 packets has mean x̄ = 998 g. Find a 95% confidence interval for μ. Is the label "1000 g" believable?

SE = 8/√64 = 1 g. Margin = 1.96 × 1 = 1.96 g. Interval: 998 ± 1.96 = (996.04, 999.96) g. 1000 g lies just outside, so the data suggest the mean is a little below 1000 g.

5. How large a sample is needed to estimate a proportion within ±4% at 95% confidence?

Use the safe value p = 0.5. n ≈ (1.96 / (2 × 0.04))² = (24.5)² = 600.25. Round up: n = 601.

6. A claim says p = 0.5. In 200 simulated samples of size 40 from p = 0.5, a share of 0.30 or less appeared 3 times. Your real sample gave 0.30. Is the claim consistent with your data?

Such a low result happened in 3/200 = 1.5% of simulations. It is very rare if the claim is true, so the data are not consistent with p = 0.5. The true proportion is probably lower.

Common mistakes

Practice quiz

1. A number that describes a whole population is called a:
2. If the sample size is made 4 times larger, the standard error becomes:
3. For 95% confidence, z* is about:
4. Which sample is most likely to be biased?
5. A 95% confidence interval for p is (0.31, 0.39). The point estimate p̂ is:

Practice: answer these yourself

Type or choose your answer, then press Check. Use a hint if you are stuck; the full solution appears after you answer.

Frequently asked questions

What is statistical inference in simple words?

It is using information from a sample to make careful statements about a whole population, together with a measure of how sure we are.

What is the difference between descriptive and inferential statistics?

Descriptive statistics only summarise the data you have (mean, graphs). Inferential statistics use that data to estimate or test something about a larger population.

Why is 1.96 used in a 95% confidence interval?

In a normal distribution, 95% of values lie within 1.96 standard deviations of the mean, so p̂ ± 1.96 SE catches the true value about 95% of the time.

Where this is taught

NetherlandsHAVO 5 (eindexamenjaar)Statistics (part 2)
NetherlandsVWO 5Statistics and probability (part 2)
Spain1º BachilleratoStochastic Sense
Spain1º BachilleratoStochastic Sense
CBSE (India)Class 12Inferential Statistics
USA (Common Core, NGSS, AP)Grade 11Inferences and conclusions from data
USA (Common Core, NGSS, AP)Grade 11Inferences and conclusions from data
Japan高校(専門学科)1〜3年Advanced Mathematics II

Learn first

Learn next

Related lessons

All Maths lessons