DNA as the genetic material
Griffith (1928), transformation: living harmless R-type Streptococcus pneumoniae + heat-killed deadly S-type → mice died, and live S bacteria were found. Something from the dead S cells 'transformed' R into S.
Avery, MacLeod and McCarty (1944): they destroyed parts of the S extract one by one. Protein-digesting enzymes (proteases) and RNase did not stop transformation; DNase did. So the transforming thing was DNA.
Hershey and Chase (1952): they grew bacteriophages with radioactive ³²P (goes into DNA) or ³⁵S (goes into protein). After infection and blending, the bacteria had ³²P but not ³⁵S. So DNA, not protein, enters the cell and carries the instructions.
What makes a good genetic material?
- Can copy itself (replicate).
- Is chemically and structurally stable.
- Can change slowly (mutate) for evolution.
- Can express itself as Mendelian traits.
RNA is less stable (its 2′-OH makes it reactive), so DNA is better for storing information; RNA is better for quick jobs. RNA was probably the first genetic material (RNA world).
Structure of DNA and RNA
A nucleotide = nitrogen base + pentose sugar + phosphate. Base + sugar = nucleoside. Purines: adenine (A), guanine (G). Pyrimidines: cytosine (C), thymine (T, only DNA), uracil (U, only RNA). DNA sugar is deoxyribose; RNA sugar is ribose.
Nucleotides join by 3′–5′ phosphodiester bonds. Watson and Crick (1953), using Rosalind Franklin and Wilkins's X-ray pictures, gave the double helix:
- Two strands, antiparallel (5′→3′ and 3′→5′).
- A=T (2 H-bonds), G≡C (3 H-bonds). A purine always faces a pyrimidine, so width stays 2 nm.
- Right-handed coil, one turn = 3.4 nm = 10 base pairs, gap between pairs 0.34 nm.
- Chargaff's rule: A = T and G = C, so (A + G) = (C + T).
RNA is usually single-stranded. Types: mRNA (message), tRNA (adaptor), rRNA (part of ribosome).
Packaging of DNA
Length of DNA = number of base pairs × 0.34 nm. Human diploid cell: 6.6 × 10⁹ bp × 0.34 × 10⁻⁹ m ≈ 2.2 m, packed into a nucleus about 10⁻⁶ m wide.
- Prokaryotes: no nucleus. Negatively charged DNA is held by positively charged proteins in a region called the nucleoid, in loops.
- Eukaryotes: DNA (negative) wraps around histones (positive, rich in lysine and arginine). Eight histones form a histone octamer; about 200 bp of DNA wrap it = one nucleosome. Nucleosomes in a row look like 'beads on a string' → chromatin fibre → coiled → chromosome at metaphase. Non-histone chromosomal (NHC) proteins help further packing.
Euchromatin: loosely packed, light-stained, active. Heterochromatin: tightly packed, dark-stained, inactive.
DNA replication (semi-conservative)
Watson and Crick predicted that each strand serves as a template. Meselson and Stahl (1958) proved it: they grew E. coli in heavy ¹⁵N, then moved them to normal ¹⁴N. After 1 generation all DNA was hybrid (¹⁵N–¹⁴N); after 2 generations half hybrid, half light. Taylor showed the same in bean root tips using radioactive thymidine.
How it happens
- Helicase opens the helix at the origin of replication, making a Y-shaped replication fork.
- DNA-dependent DNA polymerase adds nucleotides only in the 5′→3′ direction, using deoxynucleoside triphosphates (they give both the unit and the energy).
- On one strand the new DNA grows continuously (leading strand); on the other it is made in short pieces (lagging strand, Okazaki fragments).
- DNA ligase joins the pieces.
It is very fast (E. coli copies 4.6 × 10⁶ bp in about 18 minutes, roughly 2000 bp per second) and very accurate. In eukaryotes it happens in the S-phase.
Central dogma
Francis Crick proposed that genetic information flows DNA → RNA → protein. DNA → DNA is replication, DNA → RNA is transcription, RNA → protein is translation.
Reverse transcription (RNA → DNA) happens in some viruses like HIV, using reverse transcriptase. So the flow can sometimes go backwards (central dogma reverse).
Transcription
Only one strand is copied. The template strand (3′→5′) is read; the other is the coding strand (5′→3′), which has the same sequence as the RNA except T in place of U.
A transcription unit has a promoter (upstream, where RNA polymerase binds), the structural gene and a terminator.
In bacteria
One RNA polymerase makes all RNAs. With the σ (sigma) factor it starts (initiation), it grows the RNA (elongation), and with the ρ (rho) factor it stops (termination). Transcription and translation can happen together because there is no nucleus.
In eukaryotes
- RNA pol I → rRNAs (28S, 18S, 5.8S); RNA pol II → hnRNA (mRNA precursor); RNA pol III → tRNA, 5S rRNA, snRNA.
- Genes are split: exons (expressed) and introns (intervening). hnRNA is processed: splicing removes introns; capping adds methyl guanosine triphosphate at the 5′ end; tailing adds 200–300 adenines at the 3′ end (poly-A tail).
Genetic code
George Gamow reasoned: 4 bases must code 20 amino acids. 4¹ = 4 and 4² = 16 are too few; 4³ = 64 is enough. So the code is a triplet. Nirenberg, Khorana and Ochoa worked out the codons.
- 64 codons: 61 code amino acids, 3 are stop codons (UAA, UAG, UGA).
- AUG = methionine and also the start codon (dual role).
- Unambiguous and specific: one codon → one amino acid.
- Degenerate: many amino acids have more than one codon.
- Read without commas, in a continuous frame.
- Nearly universal: UUU = Phe in bacteria and humans (a few exceptions in mitochondria).
Mutations and the code
A point mutation (one base changed) can change one amino acid (sickle-cell). An insertion or deletion of 1 or 2 bases shifts the reading frame (frameshift); adding or removing 3 bases adds or removes one amino acid and keeps the frame. tRNA is the adaptor: its anticodon loop reads the codon and its 3′ end (CCA) carries the amino acid. It looks like a clover-leaf in 2D and an inverted L in 3D.
Translation
- Charging of tRNA (aminoacylation): each amino acid is attached to its tRNA, using ATP.
- Initiation: the small ribosome subunit meets the mRNA at AUG; the large subunit joins. The ribosome has two sites for tRNAs.
- Elongation: tRNAs bring amino acids codon by codon; peptide bonds form (the rRNA 23S acts as the enzyme, a ribozyme). The ribosome moves along 5′→3′.
- Termination: at a stop codon, a release factor frees the polypeptide.
The mRNA also has untranslated regions (UTRs) before AUG and after the stop codon; they help efficient translation.
Regulation of gene expression: the lac operon
Cells switch genes on only when needed. In bacteria, several genes of one job sit together with one promoter and one operator = an operon. Jacob and Monod described the lac operon of E. coli:
- i gene: makes the repressor protein (always on).
- z gene: β-galactosidase — breaks lactose into glucose and galactose.
- y gene: permease — lets lactose into the cell.
- a gene: transacetylase.
No lactose: the repressor sits on the operator and blocks RNA polymerase → genes off.
Lactose present: lactose (the inducer) binds the repressor, changes its shape, and it falls off the operator → RNA polymerase transcribes z, y, a → enzymes made. This is negative regulation of an inducible operon. When lactose is used up, the repressor binds again.
Human Genome Project and Rice Genome
The Human Genome Project (HGP) ran from 1990 to 2003 to read all ~3 × 10⁹ bp of human DNA. Methods: Expressed Sequence Tags (find expressed genes) and sequence annotation (read everything, then find genes). DNA was cut into pieces, cloned in BAC/YAC vectors, sequenced with automated machines (Sanger method), and joined using overlaps by computers.
Key findings
- About 3164.7 million bp; average gene about 3000 bases; the biggest gene, dystrophin, 2.4 million bases.
- About 30,000 genes (far fewer than thought); 99.9% of bases are the same in all people.
- Less than 2% of the genome codes for proteins; large parts are repeats.
- Chromosome 1 has the most genes (2968); the Y has the fewest (231).
- About 1.4 million places with single base differences (SNPs).
Other organisms sequenced include rice, bacteria, yeast, Caenorhabditis elegans, fruit fly and Arabidopsis. The rice genome helps find genes for yield, disease resistance and stress tolerance.
DNA fingerprinting
99.9% of DNA is the same in all humans, so we compare the 0.1% that differs, mainly repetitive DNA. When DNA is spun in density gradient centrifugation, repeats form small peaks called satellite DNA. They do not code proteins and vary a lot between people (polymorphism). Alec Jeffreys developed the method using VNTRs (Variable Number of Tandem Repeats).
- Isolate DNA (blood, hair root, saliva, semen).
- Cut it with restriction enzymes.
- Separate fragments by gel electrophoresis.
- Transfer to a nylon/nitrocellulose membrane (Southern blotting).
- Add a labelled VNTR probe that sticks to matching pieces (hybridisation).
- Detect bands on X-ray film (autoradiography).
A child gets half its bands from the mother and half from the father. Uses: forensic cases, paternity tests, population and evolution studies. PCR can multiply tiny samples first.
Key formulas and definitions
- Base pairing: A = T (2 H-bonds), G ≡ C (3 H-bonds)
- Chargaff: A = T, G = C, so A + G = C + T (purines = pyrimidines)
- Length of DNA = number of bp × 0.34 nm
- One helix turn = 10 bp = 3.4 nm
- Codons = 4³ = 64 (61 sense + 3 stop)
- Amino acids in a protein = (coding bases ÷ 3) − 1 stop codon
- Meselson–Stahl after n generations in ¹⁴N: hybrid fraction = 2/2ⁿ, light = 1 − 2/2ⁿ
Worked examples
1. A DNA sample has 20% adenine. Find the % of T, G and C.
Step 1: A = T, so T = 20%. Step 2: A + T = 40%, so G + C = 60%. Step 3: G = C = 30% each.
2. A DNA molecule has 1000 base pairs, of which 300 are A–T pairs. How many hydrogen bonds hold the two strands?
Step 1: G–C pairs = 1000 − 300 = 700. Step 2: H-bonds = 300 × 2 + 700 × 3 = 600 + 2100. Answer: 2700 H-bonds.
3. Find the length of a DNA molecule with 5386 base pairs (phage φX174 size).
Step 1: Length = bp × 0.34 nm. Step 2: 5386 × 0.34 = 1831.24 nm. Answer: about 1.83 µm.
4. Write the mRNA made from the template strand 3′-TACGCAATG-5′.
Step 1: Pair each base: T→A, A→U, C→G, G→C. Step 2: mRNA 5′-AUGCGUUAC-3′. Step 3: It equals the coding strand 5′-ATGCGTTAC-3′ with U for T.
5. An mRNA reads 5′-AUG UUU GGC AAA UGA-3′. How many amino acids will the protein have, and which?
Step 1: Read codons: AUG (Met), UUU (Phe), GGC (Gly), AAA (Lys), UGA (stop). Step 2: Stop adds no amino acid. Answer: 4 amino acids: Met–Phe–Gly–Lys.
6. A protein has 300 amino acids. What is the minimum number of bases in its coding part of mRNA (with stop codon)?
Step 1: 300 codons for amino acids + 1 stop = 301 codons. Step 2: 301 × 3 = 903 bases.
7. E. coli with only ¹⁵N DNA is moved to ¹⁴N medium. After 3 generations, what fraction of DNA molecules is hybrid?
Step 1: Molecules after 3 generations = 2³ = 8 (from one). Step 2: Only 2 keep an old ¹⁵N strand → hybrid. Step 3: 2/8 = 1/4 hybrid, 6/8 = 3/4 light.
8. In the mRNA AUG CCU AAA UGA, one base (C at position 4) is deleted. What happens?
Step 1: New sequence AUG CUA AAU GA… Step 2: Codons change after the deletion: Met–Leu–Asn… Step 3: This is a frameshift mutation: all later amino acids change, and the stop codon may be lost.
Common mistakes
- Saying the coding strand is copied during transcription. The template strand is read; the RNA matches the coding strand (with U).
- Counting a stop codon as an amino acid. Stop codons add nothing.
- Thinking the repressor of the lac operon is made only when lactose is absent. The i gene is always on; lactose only switches the repressor off.
- Using Chargaff's rule for single-stranded DNA or RNA. A = T and G = C only work for double strands.