The idea of a hypothesis test
We want to know something about a whole population (all tosses of a coin, all bottles from a machine), but we only see a sample. A hypothesis test asks: is the sample result so unusual that we should doubt the old belief?
- Null hypothesis H₀: the "nothing new" statement, with an exact value, e.g. p = 0.5 or μ = 500.
- Alternative hypothesis H₁: what we suspect, e.g. p > 0.5 (one-tailed) or p ≠ 0.5 (two-tailed).
- Test statistic: the number we compute from the sample (e.g. X = number of heads).
- Significance level α: how rare a result must be before we reject H₀, often 5%.
Critical region and p-value
The critical region is the set of values of the test statistic that make us reject H₀. The first value in it is the critical value.
The p-value is the probability, if H₀ is true, of getting the observed result or something more extreme.
Decision rule: if p-value ≤ α (or the result is in the critical region), reject H₀. Otherwise do not reject H₀. We never say we "proved" H₀; we only lacked evidence against it.
Worked idea (binomial, one-tailed)
X ~ B(20, 0.5). P(X ≥ 15) = 0.0207 ≤ 0.05 but P(X ≥ 14) = 0.0577 > 0.05. So the critical region is X ≥ 15, and the actual significance level is 2.07%.
One-tailed and two-tailed tests
If H₁ says "greater than" or "less than", all of α goes in one tail. If H₁ says "not equal", split α: for 5%, put 2.5% in each tail.
Example: H₀: p = 0.5, H₁: p ≠ 0.5, n = 20, 5%. Lower tail: P(X ≤ 5) = 0.0207 ≤ 0.025. Upper tail: P(X ≥ 15) = 0.0207. Critical region: X ≤ 5 or X ≥ 15.
Tests for a mean and for proportions (normal)
For a large sample, the sample mean x̄ is close to normal. To test H₀: μ = μ₀ with known σ, use
z = (x̄ − μ₀) ÷ (σ / √n)
Compare z with the critical value: 1.645 (one-tailed 5%), ±1.96 (two-tailed 5%), 2.326 (one-tailed 1%). If σ is unknown and n is small, use a t-test with the sample standard deviation.
For a proportion: z = (p̂ − p₀) ÷ √(p₀(1 − p₀)/n). For two means or two proportions, the numerator becomes the difference, and the standard errors are combined.
A 95% confidence interval and a two-tailed 5% test agree: H₀ is rejected exactly when μ₀ lies outside the interval.
Type I and Type II errors
- Type I error: reject H₀ when it is true. P(Type I) = actual significance level (e.g. 2.07%).
- Type II error: do not reject H₀ when it is false. Its probability depends on the true value of the parameter.
- Power = 1 − P(Type II) = chance of correctly rejecting a false H₀. Bigger samples raise power.
Making α smaller cuts Type I errors but makes Type II errors more likely.
Key formulas and definitions
- p-value = P(result as extreme or more | H₀ true)
- Reject H₀ if p-value ≤ α
- Binomial: X ~ B(n, p₀) under H₀
- z = (x̄ − μ₀) / (σ/√n)
- Proportion: z = (p̂ − p₀) / √(p₀(1 − p₀)/n)
- Power = 1 − P(Type II error)
Worked examples
1. A coin is tossed 20 times; 16 heads. Test at 5% whether it favours heads.
H₀: p = 0.5, H₁: p > 0.5, X ~ B(20, 0.5). p-value = P(X ≥ 16) = 0.0059. 0.0059 < 0.05, so reject H₀. There is evidence the coin favours heads.
2. Same test, but we got 14 heads.
P(X ≥ 14) = 0.0577 > 0.05. Do not reject H₀. There is not enough evidence that the coin is biased.
3. Find the critical region for H₀: p = 0.5, H₁: p > 0.5, n = 20 at 5%, and the probability of a Type I error.
P(X ≥ 15) = 0.0207 ≤ 0.05, P(X ≥ 14) = 0.0577 > 0.05. Critical region X ≥ 15. P(Type I) = 0.0207.
4. A machine should fill 500 ml. σ = 4 ml. A sample of 16 bottles has mean 497.5 ml. Test at 5% (two-tailed).
z = (497.5 − 500) ÷ (4/4) = −2.5. |−2.5| > 1.96, so reject H₀. The machine seems to be filling wrongly.
5. A seed company says 80% of seeds grow. In 200 seeds, 148 grow. Test at 5% whether the rate is lower.
p̂ = 0.74. z = (0.74 − 0.8) ÷ √(0.8 × 0.2 / 200) = −0.06 ÷ 0.02828 = −2.12. One-tailed critical value −1.645. −2.12 < −1.645, reject H₀: the growth rate seems lower.
6. For n = 20, H₀: p = 0.5 at 5% (critical region X ≥ 15), the true p is 0.7. Find P(Type II error).
Type II = P(X ≤ 14 | p = 0.7). From tables, P(X ≤ 14) ≈ 0.584. So there is about a 58% chance of missing this bias. Power ≈ 0.416.
Common mistakes
- Writing H₀ with an inequality (p > 0.5). H₀ always has an exact value.
- Saying "H₀ is proved true". We only say there is not enough evidence to reject it.
- Using P(X = 16) instead of P(X ≥ 16) for the p-value.
- Forgetting to halve α in a two-tailed test.