📘 CodingMarble Learn

Hypothesis Testing

A hypothesis test checks a claim about a population using a sample. Start with the null hypothesis H₀ (no change, e.g. p = 0.5) and the alternative H₁ (what we suspect, e.g. p > 0.5). Choose a significance level such as 5%. Work out how likely the sample result (or more extreme) is if H₀ were true: the p-value. If the p-value is below the level, or the result falls in the critical region, reject H₀. Otherwise there is not enough evidence to reject it. Type I error = rejecting a true H₀; Type II error = not rejecting a false H₀.

🎬 Step-by-step story

  1. A friend says: "This coin lands heads more often." We test it. H₀: p = 0.5 (the coin is fair). H₁: p > 0.5 (it favours heads).
  2. Suppose H₀ is true and we toss 20 times. These bars show how likely each number of heads is. The middle, 10, is most likely. Far right is rare.
  3. We pick a 5% significance level. The red bars on the far right add up to less than 5%. X ≥ 15 is the critical region.
  4. We toss and get 16 heads. 16 is inside the red region. The p-value P(X ≥ 16) is only 0.6%, below 5%. We reject H₀.
  5. Tests can be wrong. A fair coin can still land in the red (2.1% of the time): a Type I error. A biased coin can miss the red: a Type II error.
  6. Try it: move the observed heads and change the level to 10% or 1%. Watch the red region and the decision change.

Tip: drag the 3D scene to turn it. Use two fingers to zoom.

🤔 Common doubts, cleared

Why does H₀ get an exact value like p = 0.5?

We need one exact value to draw the probability bars and work out how rare a result is. Step 1 builds those bars from p = 0.5.

Why is the critical region only on the right?

H₁ says p > 0.5, so only many heads count as evidence. Few heads would not support 'favours heads'.

Why is the actual level 2.1% and not exactly 5%?

Heads come in whole numbers. Adding the X = 14 bar would take the tail to 5.8%, over 5%, so we stop at 15 and get 2.1%.

Why do we use P(X ≥ 16) and not P(X = 16)?

Any single value has a small chance. We ask how likely a result this extreme or more is. The p-value adds all bars from 16 upward.

If we reject H₀, can we still be wrong?

Yes. A fair coin lands in the red region about 2.1% of the time. That is a Type I error.

What changes if I choose a 1% level?

The red region shrinks to X ≥ 16, so stronger evidence is needed. Try it with the level menu.

The idea of a hypothesis test

We want to know something about a whole population (all tosses of a coin, all bottles from a machine), but we only see a sample. A hypothesis test asks: is the sample result so unusual that we should doubt the old belief?

Critical region and p-value

The critical region is the set of values of the test statistic that make us reject H₀. The first value in it is the critical value.

The p-value is the probability, if H₀ is true, of getting the observed result or something more extreme.

Decision rule: if p-value ≤ α (or the result is in the critical region), reject H₀. Otherwise do not reject H₀. We never say we "proved" H₀; we only lacked evidence against it.

Worked idea (binomial, one-tailed)

X ~ B(20, 0.5). P(X ≥ 15) = 0.0207 ≤ 0.05 but P(X ≥ 14) = 0.0577 > 0.05. So the critical region is X ≥ 15, and the actual significance level is 2.07%.

One-tailed and two-tailed tests

If H₁ says "greater than" or "less than", all of α goes in one tail. If H₁ says "not equal", split α: for 5%, put 2.5% in each tail.

Example: H₀: p = 0.5, H₁: p ≠ 0.5, n = 20, 5%. Lower tail: P(X ≤ 5) = 0.0207 ≤ 0.025. Upper tail: P(X ≥ 15) = 0.0207. Critical region: X ≤ 5 or X ≥ 15.

Tests for a mean and for proportions (normal)

For a large sample, the sample mean x̄ is close to normal. To test H₀: μ = μ₀ with known σ, use

z = (x̄ − μ₀) ÷ (σ / √n)

Compare z with the critical value: 1.645 (one-tailed 5%), ±1.96 (two-tailed 5%), 2.326 (one-tailed 1%). If σ is unknown and n is small, use a t-test with the sample standard deviation.

For a proportion: z = (p̂ − p₀) ÷ √(p₀(1 − p₀)/n). For two means or two proportions, the numerator becomes the difference, and the standard errors are combined.

A 95% confidence interval and a two-tailed 5% test agree: H₀ is rejected exactly when μ₀ lies outside the interval.

Type I and Type II errors

Making α smaller cuts Type I errors but makes Type II errors more likely.

Key formulas and definitions

Worked examples

1. A coin is tossed 20 times; 16 heads. Test at 5% whether it favours heads.

H₀: p = 0.5, H₁: p > 0.5, X ~ B(20, 0.5). p-value = P(X ≥ 16) = 0.0059. 0.0059 < 0.05, so reject H₀. There is evidence the coin favours heads.

2. Same test, but we got 14 heads.

P(X ≥ 14) = 0.0577 > 0.05. Do not reject H₀. There is not enough evidence that the coin is biased.

3. Find the critical region for H₀: p = 0.5, H₁: p > 0.5, n = 20 at 5%, and the probability of a Type I error.

P(X ≥ 15) = 0.0207 ≤ 0.05, P(X ≥ 14) = 0.0577 > 0.05. Critical region X ≥ 15. P(Type I) = 0.0207.

4. A machine should fill 500 ml. σ = 4 ml. A sample of 16 bottles has mean 497.5 ml. Test at 5% (two-tailed).

z = (497.5 − 500) ÷ (4/4) = −2.5. |−2.5| > 1.96, so reject H₀. The machine seems to be filling wrongly.

5. A seed company says 80% of seeds grow. In 200 seeds, 148 grow. Test at 5% whether the rate is lower.

p̂ = 0.74. z = (0.74 − 0.8) ÷ √(0.8 × 0.2 / 200) = −0.06 ÷ 0.02828 = −2.12. One-tailed critical value −1.645. −2.12 < −1.645, reject H₀: the growth rate seems lower.

6. For n = 20, H₀: p = 0.5 at 5% (critical region X ≥ 15), the true p is 0.7. Find P(Type II error).

Type II = P(X ≤ 14 | p = 0.7). From tables, P(X ≤ 14) ≈ 0.584. So there is about a 58% chance of missing this bias. Power ≈ 0.416.

Common mistakes

Practice quiz

1. The null hypothesis usually says:
2. If the p-value is 0.03 and α = 5%:
3. A Type I error is:
4. For a two-tailed test at 5%, each tail gets:
5. z critical value for a two-tailed 5% test:

Practice: answer these yourself

Type or choose your answer, then press Check. Use a hint if you are stuck; the full solution appears after you answer.

Frequently asked questions

What is hypothesis testing in simple words?

It is a way to use a sample to decide whether there is enough evidence against a starting claim (H₀).

What does a p-value of 0.03 mean?

If H₀ were true, there would be a 3% chance of a result as extreme as ours or more. At the 5% level we reject H₀.

What is the difference between Type I and Type II errors?

Type I: rejecting H₀ when it is true (a false alarm). Type II: not rejecting H₀ when it is false (a missed effect).

Where this is taught

NetherlandsHAVO 5 (eindexamenjaar)Statistics and probability (part 2)
NetherlandsVWO 5Probability and statistics (part 2)
CBSE (India)Class 12Inferential Statistics
England (GCSE, A level)Year 12Optional application 2 Statistics (part 1)
England (GCSE, A level)Year 12O Statistical hypothesis testing (binomial)
England (GCSE, A level)Year 13Optional application 2 Statistics (part 2)
England (GCSE, A level)Year 13O Hypothesis testing
USA (Common Core, NGSS, AP)Grade 12Inference for Categorical Data: Proportions
USA (Common Core, NGSS, AP)Grade 12Inference for Quantitative Data: Means
Japan高校1年Data analysis
Japan高校2年Statistical inference
South Korea고등학교 2학년Analysing data
Germany (Bavaria)Jahrgangsstufe 12One-sided significance test for binomial data
China高三Elective A (science/engineering)

Learn first

Related lessons

All Maths lessons