Measuring probability by frequency
Toss a coin n times and count the number of heads, h. The relative frequency of heads is f = h ÷ n.
- For small n, f jumps around a lot. 3 heads in 10 tosses (f = 0.3) is quite normal.
- For large n, f becomes stable and stays close to the probability p.
So when we cannot calculate a probability (for example, the chance a drawing pin lands point-up), we can estimate it by doing many trials: p ≈ f. The more trials, the better the estimate.
What the law of large numbers says
Bernoulli's theorem (the law of large numbers for frequencies): in n independent trials with the same probability p of success, for every ε > 0,
P(|f − p| ≥ ε) → 0 as n → ∞.
The general form is about averages. Let X₁, X₂, …, Xₙ be independent random variables with the same mean μ and variance σ². Their sample mean is X̄ = (X₁ + … + Xₙ) ÷ n. Then the chance that X̄ is more than ε away from μ goes to 0 as n grows.
Careful: the law does not say that luck "evens out". After 5 tails in a row, the next toss is still 50 : 50. The difference between heads and tails can even grow; only the proportion settles.
Chebyshev's (Bienaymé–Chebyshev) inequality
For any random variable X with mean μ and variance σ², and any a > 0:
P(|X − μ| ≥ a) ≤ σ² ÷ a².
It works for every distribution, which is why it is called a concentration inequality: it shows the values are concentrated near the mean. It is often rough, but it is always true.
Apply it to the sample mean X̄. Since E(X̄) = μ and V(X̄) = σ² ÷ n:
P(|X̄ − μ| ≥ ε) ≤ σ² ÷ (nε²).
As n grows, the right side goes to 0, which proves the law of large numbers. For frequencies, σ² = p(1 − p) ≤ 0.25, so P(|f − p| ≥ ε) ≤ p(1 − p) ÷ (nε²) ≤ 1 ÷ (4nε²).
Choosing a sample size and testing a claim
Sample size: we want the risk of an error bigger than ε to be at most α. Make σ² ÷ (nε²) ≤ α, so n ≥ σ² ÷ (αε²). For a proportion with unknown p, use σ² ≤ 0.25: n ≥ 1 ÷ (4αε²).
Sampling: a random sample from a large population behaves like repeated trials, so the sample proportion is close to the population proportion when the sample is big and fair.
Simple test: a claim says p = 0.5. If in n trials the frequency is so far from 0.5 that Chebyshev says such a gap would happen with probability below, say, 5%, we have good reason to doubt the claim.
Key formulas and definitions
- Relative frequency f = h ÷ n
- P(|X − μ| ≥ a) ≤ σ² ÷ a² (Chebyshev)
- E(X̄) = μ, V(X̄) = σ² ÷ n
- P(|X̄ − μ| ≥ ε) ≤ σ² ÷ (nε²)
- P(|f − p| ≥ ε) ≤ p(1 − p) ÷ (nε²) ≤ 1 ÷ (4nε²)
- Sample size: n ≥ σ² ÷ (αε²)
Worked examples
1. A drawing pin is dropped 400 times and lands point-up 248 times. Estimate the probability of point-up.
f = 248 ÷ 400 = 0.62. Estimate: p ≈ 0.62.
2. A fair coin is tossed 1000 times. Use Chebyshev to bound the chance that the frequency of heads is 0.1 or more away from 0.5.
p(1 − p) = 0.25, nε² = 1000 × 0.01 = 10. Bound = 0.25 ÷ 10 = 0.025. So the chance is at most 2.5%.
3. X has mean 50 and variance 16. Bound P(|X − 50| ≥ 10).
σ² ÷ a² = 16 ÷ 100 = 0.16. So the probability is at most 0.16.
4. A die is rolled 600 times. Bound the chance that the frequency of sixes differs from 1/6 by 0.05 or more.
p(1 − p) = (1/6)(5/6) = 5/36 ≈ 0.139. nε² = 600 × 0.0025 = 1.5. Bound ≈ 0.139 ÷ 1.5 ≈ 0.093.
5. How many people must a survey ask so that the chance of an error of 0.05 or more in a proportion is at most 0.1?
n ≥ 1 ÷ (4αε²) = 1 ÷ (4 × 0.1 × 0.0025) = 1 ÷ 0.001 = 1000 people.
6. A coin is claimed to be fair. In 2000 tosses it shows 1200 heads. Is the claim believable?
f = 0.6, a gap of 0.1. Chebyshev: P(|f − 0.5| ≥ 0.1) ≤ 0.25 ÷ (2000 × 0.01) = 0.0125. A gap this big would happen at most 1.25% of the time for a fair coin, so we doubt the claim.
Common mistakes
- The gambler's fallacy: thinking tails is "due" after many heads. Each toss is independent; only the long-run proportion settles.
- Thinking the number of heads gets close to n/2. It is the proportion h/n that gets close to 1/2; the difference h − n/2 can grow.
- Forgetting to square ε (or a) in Chebyshev's inequality.
- Reading Chebyshev's bound as the exact probability. It is only an upper limit; the true chance is usually much smaller.