Statistic, parameter and sampling variability
A parameter is a number about the whole population, like the true mean μ or the true proportion p. We usually do not know it.
A statistic is a number we work out from a sample, like the sample mean x̄ or the sample proportion p̂. We use it to guess the parameter.
Different samples give different statistics. This is called sampling variability. It is not a mistake. It is just chance.
What is a sampling distribution?
Imagine taking every possible sample of size n and working out x̄ each time. The pattern of all those x̄ values is the sampling distribution of x̄. In the 3D scene, each orange block is one sample mean and the pile is the sampling distribution.
Every sampling distribution is described by three things: centre, spread and shape.
Sampling distribution of a sample mean
- Centre: the mean of all x̄ values is μ. So x̄ is an unbiased estimator: on average it hits the target.
- Spread: the standard deviation of x̄ is σ/√n. This is called the standard error. Use it only when the sample is less than 10% of the population (the 10% condition), so the picks are nearly independent.
- Shape: if the population is normal, x̄ is normal for any n. If not, use the Central Limit Theorem.
Why does a bigger sample help?
Because √n is on the bottom. To halve the spread you need 4 times the sample size. To make it a third, 9 times.
The Central Limit Theorem (CLT)
The Central Limit Theorem says: when the sample size n is large, the sampling distribution of x̄ is close to normal, whatever the shape of the population. A common rule of thumb is n ≥ 30. A very skewed population needs a bigger n; a nearly symmetric one needs less.
This is why the normal curve appears everywhere in statistics. We can then use z-scores: z = (x̄ − μ) ÷ (σ/√n).
Worked idea
Waiting times: μ = 5 min, σ = 5 min, n = 40. Standard error = 5/√40 ≈ 0.79 min. What is the chance a sample mean is more than 6.5 min? z = (6.5 − 5)/0.79 ≈ 1.90, so P ≈ 0.029, about 3%.
Sampling distribution of a sample proportion
When each person is a yes/no answer, we use the sample proportion p̂ = (number of yes) ÷ n.
- Centre: p (p̂ is unbiased).
- Spread: √(p(1−p)/n).
- Shape: close to normal when the large counts condition holds: np ≥ 10 and n(1−p) ≥ 10.
Example: 30% of phones in a city use a cracked screen protector, n = 100. SE = √(0.3 × 0.7/100) ≈ 0.046. np = 30 and n(1−p) = 70, both ≥ 10, so a normal model is fine.
Differences between two samples
Often we compare two groups, like two schools or two medicines. Take independent samples from each group.
Two proportions: p̂₁ − p̂₂
Centre = p₁ − p₂. Spread = √(p₁(1−p₁)/n₁ + p₂(1−p₂)/n₂). Normal if all four counts (n₁p₁, n₁(1−p₁), n₂p₂, n₂(1−p₂)) are at least 10.
Two means: x̄₁ − x̄₂
Centre = μ₁ − μ₂. Spread = √(σ₁²/n₁ + σ₂²/n₂). Normal if both populations are normal or both samples are large (CLT).
Key idea: variances add, even when you subtract. So the difference is always more spread out than each part.
Quality assurance: control charts
Factories use sampling distributions every day. They take a small sample at regular times and plot its mean on a control chart.
- A centre line is drawn at the target mean.
- Warning limits are drawn at about target ± 2 standard errors. About 95% of sample means fall inside.
- Action limits are at about target ± 3 standard errors. About 99.8% fall inside.
One point outside the action limits, or two points in a row outside the warning limits, means: stop and check the machine.
Try it: the dice average game
Roll one die 20 times and write the results: the pattern is flat (each face about equal). Now roll 4 dice at once, find their mean, and do this 20 times. Plot both on dot plots. The means crowd near 3.5 and look like a hill. That hill is a sampling distribution, and the change of shape is the CLT at work. Then try it in the 3D: set n = 1, then n = 4, then n = 30.
Key formulas and definitions
- Mean of x̄ = μ; SD of x̄ (standard error) = σ/√n
- Mean of p̂ = p; SD of p̂ = √(p(1−p)/n)
- z = (x̄ − μ) ÷ (σ/√n)
- x̄₁ − x̄₂: centre μ₁ − μ₂, SD = √(σ₁²/n₁ + σ₂²/n₂)
- p̂₁ − p̂₂: centre p₁ − p₂, SD = √(p₁(1−p₁)/n₁ + p₂(1−p₂)/n₂)
- Conditions: random sample; n ≤ 10% of population; normal shape (CLT n ≥ 30, or np ≥ 10 and n(1−p) ≥ 10)
Worked examples
1. Heights of a large group of 15-year-olds: μ = 160 cm, σ = 8 cm. Samples of n = 16 are taken. Find the mean and standard error of x̄.
Mean of x̄ = 160 cm. SE = 8/√16 = 8/4 = 2 cm.
2. Same group. What sample size gives a standard error of 1 cm?
8/√n = 1 → √n = 8 → n = 64. (Halving the SE needs 4 times the sample: 16 → 64.)
3. Bottles are filled with μ = 500 mL, σ = 4 mL. A sample of 36 bottles is taken. Find P(x̄ < 498.5 mL).
SE = 4/√36 = 0.667 mL. z = (498.5 − 500)/0.667 = −2.25. P(z < −2.25) ≈ 0.012, about 1.2%.
4. In a large town, 40% of homes have solar panels. A random sample of 150 homes is taken. Describe the sampling distribution of p̂.
Centre 0.40. SE = √(0.4 × 0.6/150) = √0.0016 = 0.04. np = 60 ≥ 10 and n(1−p) = 90 ≥ 10, so it is approximately normal: N(0.40, 0.04).
5. Using the last example, find P(p̂ > 0.48).
z = (0.48 − 0.40)/0.04 = 2. P(z > 2) ≈ 0.023, about 2.3%. Getting 48% or more would be quite unusual.
6. School A: p₁ = 0.6 of students cycle, n₁ = 100. School B: p₂ = 0.5, n₂ = 100. Find the mean and SD of p̂₁ − p̂₂.
Centre = 0.6 − 0.5 = 0.1. SD = √(0.6·0.4/100 + 0.5·0.5/100) = √(0.0024 + 0.0025) = √0.0049 = 0.07.
7. Machine target 250 g, σ = 6 g, samples of n = 9. Find the warning and action limits for the control chart of the means.
SE = 6/√9 = 2 g. Warning limits: 250 ± 2 × 2 = 246 g and 254 g. Action limits: 250 ± 3 × 2 = 244 g and 256 g.
Common mistakes
- Using σ instead of σ/√n for a sample mean. The spread of the means is smaller than the spread of single values.
- Thinking the CLT makes the population normal. It only makes the distribution of the sample mean normal; the population stays skewed.
- Subtracting variances for a difference. For x̄₁ − x̄₂ you add: √(σ₁²/n₁ + σ₂²/n₂).
- Skipping the conditions: check random sampling, the 10% condition and the normal shape (n ≥ 30, or np ≥ 10 and n(1−p) ≥ 10) before using z.