📘 CodingMarble Learn

Investigating Diversity: Measuring Genetic Diversity, Sampling and Standard Deviation

Genetic diversity within or between species can be measured by comparing observable characteristics, DNA base sequences, mRNA base sequences or amino acid sequences: the fewer the differences, the closer the relationship. Since we cannot measure every organism, we take a random sample large enough to be representative, then describe it with a mean and a standard deviation (SD). SD shows the spread around the mean; about 68% of normally distributed values lie within ±1 SD. Overlapping ±SD ranges suggest a difference may not be real.

🎬 Step-by-step story

  1. Here is a field of plants. Each plant has a different height. This is variation.
  2. We cannot measure every plant. So we pick plants using random numbers. A random sample avoids bias.
  3. We group the sample heights into a chart. The red line is the mean: add all values and divide by how many.
  4. Standard deviation shows how spread out the values are. The blue band is ±1 SD. About two thirds of values sit inside it.
  5. We can also compare DNA. Line up the bases of three species. Fewer differences means a closer relationship.
  6. Free play: change the sample size and the spread. Take a new sample and watch the mean and SD change.

Tip: drag the 3D scene to turn it. Use two fingers to zoom.

🤔 Common doubts, cleared

Why do plants of the same species differ in height?

Different alleles and different conditions (light, water, soil) both cause variation.

Why not just choose plants that look typical?

Your choice would be biased; random coordinates give every plant the same chance.

If two samples have the same mean, are they the same?

No. One may be tightly grouped and the other very spread out; the SD shows this.

Why divide by n − 1 and not n?

A sample tends to underestimate the spread of the whole population; dividing by n − 1 corrects for that.

How do DNA differences show relatedness?

Mutations build up over time, so species that split recently have had less time to collect differences.

Why does the mean jump around with small samples?

With few values, one unusual plant has a big effect. Increase n in free play and it settles.

Ways to measure genetic diversity

Genetic diversity is the number of different alleles of genes in a population. We can investigate it, and how related species are, in four main ways:

Modern methods use gene technology to read DNA directly, which is far more precise than comparing visible features. The fewer the differences between two sequences, the more recently the two groups shared a common ancestor.

Random sampling: getting fair data

We usually study a sample, not the whole population. A good sample is:

Chance can never be removed, so we use statistics to judge whether a difference is likely to be real.

Mean and standard deviation

Mean x̄ = Σx ÷ n.

Standard deviation s = √( Σ(x − x̄)² ÷ (n − 1) ). Steps: find the mean; subtract it from each value; square each difference; add them; divide by n − 1; take the square root.

When two samples are shown as mean ± SD bars: if the bars overlap a lot, the difference between means may be due to chance; if they do not overlap, the difference is more likely to be real (a statistical test confirms it).

Exam focus

Expect to: calculate a mean and SD from a small table; interpret error bars; explain why sampling must be random and large; explain why comparing DNA is better than comparing visible traits; and use sequence differences to say which species are most closely related.

Key formulas and definitions

Worked examples

1. Five leaf lengths (mm): 40, 42, 44, 46, 48. Find the mean and SD.

Mean = 220 ÷ 5 = 44 mm. Differences: −4, −2, 0, 2, 4. Squares: 16, 4, 0, 4, 16 → sum 40. Divide by n − 1 = 4 → 10. SD = √10 ≈ 3.16 mm.

2. A gene section has 12 bases. Species B differs from A at 1 base; species C differs from A at 5 bases. Which is more closely related to A? Give % differences.

B: 1/12 × 100 ≈ 8.3%. C: 5/12 × 100 ≈ 41.7%. B is more closely related to A because fewer bases differ.

3. Two populations of snails have shell widths 18 ± 1.5 mm and 20 ± 4 mm (mean ± SD). Comment.

The ranges 16.5–19.5 and 16–24 overlap a lot, so the difference in means may be due to chance. The second population is much more varied (larger SD).

4. Why may two species have identical amino acid sequences for a protein but different DNA?

The code is degenerate: several triplets code for the same amino acid. A base change in the third position often gives the same amino acid, so DNA differs but the protein does not.

5. Values 3, 5, 7. Find the SD.

Mean = 5. Differences −2, 0, 2; squares 4, 0, 4; sum 8; ÷ (3 − 1) = 4; SD = √4 = 2.

Common mistakes

Practice quiz

1. Which method gives the most precise measure of genetic diversity?
2. Why is sampling done at random?
3. A small standard deviation shows:
4. About what % of normally distributed values lie within ±1 SD?
5. Two species differ at 2 of 20 DNA bases. % difference:

Practice: answer these yourself

Type or choose your answer, then press Check. Use a hint if you are stuck; the full solution appears after you answer.

Frequently asked questions

How is genetic diversity measured?

By comparing observable traits, DNA base sequences, mRNA base sequences or amino acid sequences; DNA sequencing is the most precise.

What does standard deviation show in biology?

How spread out the data are around the mean. A larger SD means more variation.

Why is random sampling important?

It avoids bias, so the sample fairly represents the population.

Where this is taught

England (GCSE, A level)Year 123.4 Genetic information, variation and relationships

Learn first

Related lessons

All Biology lessons