📘 CodingMarble Learn

Sampling Distributions and the Central Limit Theorem

A statistic (like a sample mean x̄ or a sample proportion p̂) changes from sample to sample. If you took every possible sample and plotted the statistic, you would get its sampling distribution. Its centre is the true population value (μ or p), its spread is the standard error (σ/√n for means, √(p(1−p)/n) for proportions), and by the Central Limit Theorem its shape becomes close to normal when n is large enough.

🎬 Step-by-step story

  1. The population: the waiting times of thousands of people. The shape is skewed. The mean is μ = 5 minutes.
  2. Take one sample of 10 people. Its mean x̄ is close to 5, but not exactly 5.
  3. Take 400 samples. Pile up all the sample means. This pile is the sampling distribution. Its centre is μ.
  4. Make each sample bigger (n = 40). The pile gets thinner. Spread = σ/√n.
  5. Central Limit Theorem: even from a skewed population, the pile of means becomes bell-shaped when n is big.
  6. Your turn: change n, switch between mean and proportion, and check when the normal model is OK.

Tip: drag the 3D scene to turn it. Use two fingers to zoom.

🤔 Common doubts, cleared

Is the sampling distribution the same as the distribution of my sample?

No. Your sample is a handful of values (step 2). The sampling distribution is the pattern of a statistic over many, many samples (step 3).

Why is the spread of means smaller than the spread of the data?

In a sample, high and low values cancel each other. So averages bunch near the centre. The bigger the sample, the more they cancel.

Does the CLT make my data normal?

No. In step 5 the blue population stays skewed. Only the orange pile of means turns bell-shaped.

Why √n and not n?

Variance of a mean is σ²/n. Standard deviation is the square root, so it is σ/√n. That is why you need 4 times the data to halve the error.

When is a sample too small for the normal model of p̂?

When np or n(1−p) is below 10. Switch to proportion mode, set p = 0.1 and n = 30, and the check turns red.

Statistic, parameter and sampling variability

A parameter is a number about the whole population, like the true mean μ or the true proportion p. We usually do not know it.

A statistic is a number we work out from a sample, like the sample mean x̄ or the sample proportion p̂. We use it to guess the parameter.

Different samples give different statistics. This is called sampling variability. It is not a mistake. It is just chance.

What is a sampling distribution?

Imagine taking every possible sample of size n and working out x̄ each time. The pattern of all those x̄ values is the sampling distribution of x̄. In the 3D scene, each orange block is one sample mean and the pile is the sampling distribution.

Every sampling distribution is described by three things: centre, spread and shape.

Sampling distribution of a sample mean

Why does a bigger sample help?

Because √n is on the bottom. To halve the spread you need 4 times the sample size. To make it a third, 9 times.

The Central Limit Theorem (CLT)

The Central Limit Theorem says: when the sample size n is large, the sampling distribution of x̄ is close to normal, whatever the shape of the population. A common rule of thumb is n ≥ 30. A very skewed population needs a bigger n; a nearly symmetric one needs less.

This is why the normal curve appears everywhere in statistics. We can then use z-scores: z = (x̄ − μ) ÷ (σ/√n).

Worked idea

Waiting times: μ = 5 min, σ = 5 min, n = 40. Standard error = 5/√40 ≈ 0.79 min. What is the chance a sample mean is more than 6.5 min? z = (6.5 − 5)/0.79 ≈ 1.90, so P ≈ 0.029, about 3%.

Sampling distribution of a sample proportion

When each person is a yes/no answer, we use the sample proportion p̂ = (number of yes) ÷ n.

Example: 30% of phones in a city use a cracked screen protector, n = 100. SE = √(0.3 × 0.7/100) ≈ 0.046. np = 30 and n(1−p) = 70, both ≥ 10, so a normal model is fine.

Differences between two samples

Often we compare two groups, like two schools or two medicines. Take independent samples from each group.

Two proportions: p̂₁ − p̂₂

Centre = p₁ − p₂. Spread = √(p₁(1−p₁)/n₁ + p₂(1−p₂)/n₂). Normal if all four counts (n₁p₁, n₁(1−p₁), n₂p₂, n₂(1−p₂)) are at least 10.

Two means: x̄₁ − x̄₂

Centre = μ₁ − μ₂. Spread = √(σ₁²/n₁ + σ₂²/n₂). Normal if both populations are normal or both samples are large (CLT).

Key idea: variances add, even when you subtract. So the difference is always more spread out than each part.

Quality assurance: control charts

Factories use sampling distributions every day. They take a small sample at regular times and plot its mean on a control chart.

One point outside the action limits, or two points in a row outside the warning limits, means: stop and check the machine.

Try it: the dice average game

Roll one die 20 times and write the results: the pattern is flat (each face about equal). Now roll 4 dice at once, find their mean, and do this 20 times. Plot both on dot plots. The means crowd near 3.5 and look like a hill. That hill is a sampling distribution, and the change of shape is the CLT at work. Then try it in the 3D: set n = 1, then n = 4, then n = 30.

Key formulas and definitions

Worked examples

1. Heights of a large group of 15-year-olds: μ = 160 cm, σ = 8 cm. Samples of n = 16 are taken. Find the mean and standard error of x̄.

Mean of x̄ = 160 cm. SE = 8/√16 = 8/4 = 2 cm.

2. Same group. What sample size gives a standard error of 1 cm?

8/√n = 1 → √n = 8 → n = 64. (Halving the SE needs 4 times the sample: 16 → 64.)

3. Bottles are filled with μ = 500 mL, σ = 4 mL. A sample of 36 bottles is taken. Find P(x̄ < 498.5 mL).

SE = 4/√36 = 0.667 mL. z = (498.5 − 500)/0.667 = −2.25. P(z < −2.25) ≈ 0.012, about 1.2%.

4. In a large town, 40% of homes have solar panels. A random sample of 150 homes is taken. Describe the sampling distribution of p̂.

Centre 0.40. SE = √(0.4 × 0.6/150) = √0.0016 = 0.04. np = 60 ≥ 10 and n(1−p) = 90 ≥ 10, so it is approximately normal: N(0.40, 0.04).

5. Using the last example, find P(p̂ > 0.48).

z = (0.48 − 0.40)/0.04 = 2. P(z > 2) ≈ 0.023, about 2.3%. Getting 48% or more would be quite unusual.

6. School A: p₁ = 0.6 of students cycle, n₁ = 100. School B: p₂ = 0.5, n₂ = 100. Find the mean and SD of p̂₁ − p̂₂.

Centre = 0.6 − 0.5 = 0.1. SD = √(0.6·0.4/100 + 0.5·0.5/100) = √(0.0024 + 0.0025) = √0.0049 = 0.07.

7. Machine target 250 g, σ = 6 g, samples of n = 9. Find the warning and action limits for the control chart of the means.

SE = 6/√9 = 2 g. Warning limits: 250 ± 2 × 2 = 246 g and 254 g. Action limits: 250 ± 3 × 2 = 244 g and 256 g.

Common mistakes

Practice quiz

1. The centre of the sampling distribution of x̄ is:
2. If n is made 4 times bigger, the standard error of x̄:
3. The Central Limit Theorem says that for large n:
4. Standard error of p̂ when p = 0.5 and n = 100:
5. For p̂ to be approximately normal we need:

Practice: answer these yourself

Type or choose your answer, then press Check. Use a hint if you are stuck; the full solution appears after you answer.

Frequently asked questions

What is a sampling distribution in simple words?

It is the pattern you would get if you took many samples of the same size and plotted a statistic, like the mean, from each one.

What is the difference between standard deviation and standard error?

Standard deviation measures spread of individual values. Standard error measures spread of a statistic, like x̄, from sample to sample: SE = σ/√n.

Why is n ≥ 30 used for the Central Limit Theorem?

It is a rule of thumb. For most populations that are not extremely skewed, sample means of 30 or more are close enough to normal.

Where this is taught

England (GCSE, A level)Year 112. Processing, representing and analysing data (part 2)
USA (Common Core, NGSS, AP)Grade 12Probability, Random Variables, and Probability Distributions
USA (Common Core, NGSS, AP)Grade 12Inference for Categorical Data: Proportions
USA (Common Core, NGSS, AP)Grade 12Inference for Quantitative Data: Means
South Korea고등학교 2학년Statistics
South Korea고등학교 3학년Statistics

Learn first

Learn next

Related lessons

All Maths lessons