Bernoulli trials: one yes/no experiment
A Bernoulli trial is an experiment with only two results. We call one result success and the other failure. Success has probability p. Failure has probability q = 1 − p.
Examples: a coin shows heads or not; a seed sprouts or not; a penalty is scored or not.
A random variable that is 1 for success and 0 for failure follows the Bernoulli law. Its mean is p and its variance is p(1 − p).
When is a situation binomial?
Check four things (remember BINS):
- Binary: each trial is success or failure.
- Independent: one trial does not change the next.
- Number fixed: we know n before we start.
- Same p: the chance of success is the same every time.
If all four hold, X = number of successes follows B(n, p). Drawing cards with replacement is binomial. Drawing without replacement is not, because p changes (see the hypergeometric section).
The binomial formula, step by step
Think of a tree diagram with n levels. One path with k successes and n − k failures has probability pᵏ qⁿ⁻ᵏ (multiply along the path, because the trials are independent).
How many such paths are there? We must choose which k of the n places are successes: ⁿCₖ ways.
P(X = k) = ⁿCₖ pᵏ qⁿ⁻ᵏ, k = 0, 1, …, n
Where ⁿCₖ = n! ÷ (k! (n − k)!). The name comes from the binomial theorem: the terms of (q + p)ⁿ are exactly these probabilities, so they add up to (q + p)ⁿ = 1ⁿ = 1.
Mean, variance and the shape of the graph
X is the sum of n Bernoulli variables, each with mean p and variance pq. So:
- Mean E(X) = np
- Variance Var(X) = npq
- Standard deviation σ = √(npq)
The tallest bar (the mode) is always near np. When p = 0.5 the graph is symmetric. When p is small the graph leans to the left; when p is big it leans to the right. When n is large, the bar chart looks like a bell curve.
Simulation, fluctuation and a first idea of testing
A computer can simulate a trial: pick a random number between 0 and 1; if it is less than p, count a success. Repeat to get samples.
Different samples give different results. This is sampling fluctuation. For a sample of size n, the share of successes f usually lies in the interval p − 1/√n to p + 1/√n (about 95% of samples, when n ≥ 25 and 0.2 ≤ p ≤ 0.8).
Idea of a test: a firm says 50% of people like its drink. In a fair sample of 100 people only 35 do. The interval is 0.5 ± 0.1 = [0.4, 0.6]. 0.35 is outside, so we doubt the claim. A poll can also be biased if the sample is not chosen at random, for example asking only people inside the firm's own shop.
Without replacement: the hypergeometric distribution
A box has N items, K of them “good”. We pick n items without putting them back. X = number of good items picked.
P(X = k) = ᴷCₖ × ᴺ⁻ᴷCₙ₋ₖ ÷ ᴺCₙ
Its mean is also n × K/N. The trials are not independent, so it is not binomial. If N is very large compared with n, the hypergeometric is almost the same as B(n, K/N).
Try it: predict, then check
Toss 4 coins together 32 times and count heads each time. First predict: B(4, 0.5) says 0 heads about 2 times, 1 head about 8 times, 2 heads about 12 times, 3 heads about 8 times, 4 heads about 2 times. Now toss and tally. Then press “Run 100 times” in the 3D with n = 4, p = 0.5 and compare.
Key formulas and definitions
- P(X = k) = ⁿCₖ pᵏ qⁿ⁻ᵏ, q = 1 − p
- ⁿCₖ = n! / (k! (n − k)!)
- E(X) = np, Var(X) = npq, σ = √(npq)
- P(X ≥ 1) = 1 − qⁿ
- Bernoulli: E = p, Var = pq
- Hypergeometric: P(X = k) = ᴷCₖ · ᴺ⁻ᴷCₙ₋ₖ / ᴺCₙ
- Fluctuation interval (≈95%): [p − 1/√n, p + 1/√n]
Worked examples
1. A fair coin is tossed 5 times. Find the probability of exactly 3 heads.
n = 5, p = 0.5, k = 3. P = ⁵C₃ × 0.5³ × 0.5² = 10 × 1/32 = 10/32 = 0.3125.
2. A bowler hits the stumps with probability 0.3 on each ball. In 6 balls, find P(exactly 2 hits).
n = 6, p = 0.3, q = 0.7. P = ⁶C₂ × 0.3² × 0.7⁴ = 15 × 0.09 × 0.2401 ≈ 0.324.
3. 5% of bulbs are faulty. A sample of 20 is chosen at random. Find P(no faulty bulb) and P(at least one).
P(X = 0) = 0.95²⁰ ≈ 0.358. P(X ≥ 1) = 1 − 0.358 = 0.642.
4. X ~ B(10, 0.4). Find the mean, variance and standard deviation.
Mean = np = 4. Variance = npq = 10 × 0.4 × 0.6 = 2.4. σ = √2.4 ≈ 1.55.
5. The mean of a binomial variable is 6 and its variance is 4.2. Find n and p.
npq ÷ np = q = 4.2 ÷ 6 = 0.7, so p = 0.3. n = 6 ÷ 0.3 = 20.
6. How many times must a fair die be rolled so that P(at least one six) > 0.9?
P(at least one six) = 1 − (5/6)ⁿ > 0.9 ⇒ (5/6)ⁿ < 0.1. (5/6)¹² ≈ 0.112, (5/6)¹³ ≈ 0.093. So n = 13 rolls.
7. A bag has 6 red and 4 blue marbles. 3 are drawn without replacement. Find P(exactly 2 red).
Hypergeometric: ⁶C₂ × ⁴C₁ ÷ ¹⁰C₃ = 15 × 4 ÷ 120 = 0.5. (With replacement it would be ³C₂ × 0.6² × 0.4 = 0.432.)
Common mistakes
- Forgetting the ⁿCₖ factor and giving only pᵏ qⁿ⁻ᵏ, which is the chance of one particular order.
- Using the binomial when drawing without replacement; then p changes and the hypergeometric is needed.
- Working out “at least one” by adding many terms instead of 1 − P(X = 0).
- Mixing up variance and standard deviation: Var = npq, σ = √(npq).