Ways to measure genetic diversity
Genetic diversity is the number of different alleles of genes in a population. We can investigate it, and how related species are, in four main ways:
- Observable characteristics (height, leaf shape, colour). Quick and cheap, but many traits are coded by several genes and are changed by the environment, so they can mislead.
- DNA base sequences. Line up the same gene from two organisms and count differences.
- mRNA base sequences. mRNA is copied from DNA, so its differences show differences in expressed genes.
- Amino acid sequences of a shared protein (for example cytochrome c). The genetic code is degenerate, so a DNA change may not change the amino acid; amino acid comparison can hide some differences.
Modern methods use gene technology to read DNA directly, which is far more precise than comparing visible features. The fewer the differences between two sequences, the more recently the two groups shared a common ancestor.
Random sampling: getting fair data
We usually study a sample, not the whole population. A good sample is:
- Random: use a random number generator to choose coordinates on a grid, so the person sampling cannot pick 'nice' plants. This avoids sampling bias.
- Large: the bigger the sample, the less a few unusual individuals change the result. This reduces the effect of chance.
Chance can never be removed, so we use statistics to judge whether a difference is likely to be real.
Mean and standard deviation
Mean x̄ = Σx ÷ n.
Standard deviation s = √( Σ(x − x̄)² ÷ (n − 1) ). Steps: find the mean; subtract it from each value; square each difference; add them; divide by n − 1; take the square root.
- A small SD means values sit close to the mean (little variation).
- A large SD means values are widely spread.
- For a normal distribution, about 68% of values lie within ±1 SD and about 95% within ±2 SD.
When two samples are shown as mean ± SD bars: if the bars overlap a lot, the difference between means may be due to chance; if they do not overlap, the difference is more likely to be real (a statistical test confirms it).
Exam focus
Expect to: calculate a mean and SD from a small table; interpret error bars; explain why sampling must be random and large; explain why comparing DNA is better than comparing visible traits; and use sequence differences to say which species are most closely related.
Key formulas and definitions
- Mean: x̄ = Σx ÷ n
- Standard deviation: s = √( Σ(x − x̄)² ÷ (n − 1) )
- ≈ 68% of values within x̄ ± 1s; ≈ 95% within x̄ ± 2s
- % difference in a sequence = (number of differences ÷ total positions) × 100
Worked examples
1. Five leaf lengths (mm): 40, 42, 44, 46, 48. Find the mean and SD.
Mean = 220 ÷ 5 = 44 mm. Differences: −4, −2, 0, 2, 4. Squares: 16, 4, 0, 4, 16 → sum 40. Divide by n − 1 = 4 → 10. SD = √10 ≈ 3.16 mm.
2. A gene section has 12 bases. Species B differs from A at 1 base; species C differs from A at 5 bases. Which is more closely related to A? Give % differences.
B: 1/12 × 100 ≈ 8.3%. C: 5/12 × 100 ≈ 41.7%. B is more closely related to A because fewer bases differ.
3. Two populations of snails have shell widths 18 ± 1.5 mm and 20 ± 4 mm (mean ± SD). Comment.
The ranges 16.5–19.5 and 16–24 overlap a lot, so the difference in means may be due to chance. The second population is much more varied (larger SD).
4. Why may two species have identical amino acid sequences for a protein but different DNA?
The code is degenerate: several triplets code for the same amino acid. A base change in the third position often gives the same amino acid, so DNA differs but the protein does not.
5. Values 3, 5, 7. Find the SD.
Mean = 5. Differences −2, 0, 2; squares 4, 0, 4; sum 8; ÷ (3 − 1) = 4; SD = √4 = 2.
Common mistakes
- Dividing by n instead of n − 1 when finding the SD of a sample.
- Thinking a big mean means big variation. Spread is shown by SD, not by the mean.
- Saying overlapping SD bars prove there is no difference. They only suggest the difference may be due to chance; a statistical test decides.
- Choosing 'typical' organisms by eye. That causes bias; use random coordinates.