What is bioinformatics?
Bioinformatics = biology + computer science + statistics. Living things hold huge amounts of information. One human genome has about 3.2 billion DNA letters. No person can read that by hand, so we use computers.
Bioinformatics helps to: store data safely, search it fast, compare sequences, find genes, and predict what a protein does.
Spreadsheets for lab data
A spreadsheet is a grid of cells in rows and columns. Each row can be one sample; each column one measurement (for example absorbance or colony count).
- Formulas calculate for you:
=AVERAGE(B2:B7)gives the mean. - Charts show trends, such as a growth curve.
- Good habits: one header row, units in the header (mg/L), no merged cells, keep raw data and results separate, save a dated copy.
Sequence databases
A DNA sequence is written with 4 letters: A, T, G, C. A protein sequence uses 20 letters, one per amino acid.
Public databases hold millions of sequences that scientists share freely. Examples: GenBank (USA), ENA (Europe) and DDBJ (Japan) share the same data every day. UniProt holds proteins; PDB holds 3D protein shapes.
Each record has an ID number, the organism, a description and the sequence. A common text format is FASTA: a line starting with > for the name, then the letters.
Sequence alignment
Alignment lines up two sequences to see where they match. The computer slides one sequence along the other and scores each position: a match scores +, a mismatch or gap scores −. The best score wins.
% identity = matching letters ÷ letters compared × 100. High identity often means the genes have a common ancestor.
Tools such as BLAST search a whole database for sequences that look like yours in seconds. Uses: naming an unknown bacterium, finding a gene in another species, tracing a virus.
Genomes, SNPs and proteins
Genomics
A genome is all the DNA of an organism. Programs scan it for start and stop signals to predict genes and count them. Humans have about 20 000 protein-coding genes.
SNPs
A SNP (single nucleotide polymorphism, say "snip") is a place where people differ by one letter. Most SNPs are harmless; some affect disease risk or how a medicine works.
Comparative and functional genomics
Comparing genomes of species shows what is shared and what is new. Functional genomics asks which genes are switched on, where and when.
Proteomics
Proteomics studies all the proteins of a cell. Software predicts a protein's shape from its letters, which helps design medicines.
Digital images in the lab
A camera on a microscope makes a digital image: a grid of tiny squares called pixels. Each pixel stores numbers for brightness or colour.
Image software can count cells, measure size or area, and compare colours. Always add a scale bar (for example 10 µm) and never edit an image in a way that changes the result.
Ethics and digital data
- Privacy: a person's DNA can reveal health and family links. Data must be anonymised and stored securely.
- Consent: people must agree, knowing how their data will be used.
- Fairness: data should not be used to deny jobs or insurance.
- Honesty: share raw data, cite the source, never fake or hide results.
Try it
Write two short words of DNA on paper strips, for example ATGCCTA and ATGACTA. Slide one under the other and count matches at each position. Work out the % identity. Then do the same in the 3D free play.
Key formulas and definitions
- % identity = matches ÷ letters compared × 100
- Mean = sum of values ÷ number of values
- DNA letters: A, T, G, C (A pairs with T, G pairs with C)
- SNP = one-letter difference at the same position
- Pixel area × number of pixels = object area
Worked examples
1. Absorbance values are 0.42, 0.55, 0.38, 0.61, 0.47 and 0.53. Find the mean.
Sum = 2.96. Mean = 2.96 ÷ 6 ≈ 0.49.
2. Align ATGCCTAG with ATGACTAG. Find % identity.
Compare 8 positions: only position 4 differs (C vs A). Matches = 7. % identity = 7 ÷ 8 × 100 = 87.5%.
3. A cell covers 22 pixels. Each pixel is 0.5 µm × 0.5 µm. Find the cell area.
One pixel = 0.25 µm². Area = 22 × 0.25 = 5.5 µm².
4. Two people's sequences: GATTACA and GATCACA. How many SNPs, and where?
Position 4: T vs C. One SNP.
Common mistakes
- Thinking bioinformatics is only coding. It starts with a biological question; the computer is the tool.
- Calculating % identity with the longer sequence length instead of the number of letters actually compared.
- Believing every SNP causes disease. Most SNPs have no effect.
- Typing data with mixed units or merged cells, which breaks spreadsheet formulas.