What is computer architecture?
Computer architecture is the plan of a computer: which parts it has and how they talk to each other. We can split (decompose) a computer into layers:
- Hardware: CPU, memory, storage, input/output devices.
- Operating system: shares the hardware between programs.
- Applications: the apps you use.
Each layer hides the details of the layer below it. This is called abstraction. In this lesson we look inside the hardware layer.
The von Neumann model
In 1945 John von Neumann described a design that almost every computer still uses. Its big idea is the stored program: the program's instructions and the data are both kept as binary numbers in the same main memory.
Main parts
- CPU (Central Processing Unit): runs instructions.
- Main memory (RAM): numbered cells; each cell has an address.
- Input and output devices.
Buses
- Address bus: carries the address of the cell to use (CPU → memory only).
- Data bus: carries the data or instruction (both ways).
- Control bus: carries signals like "read", "write" and the clock.
Because instructions and data share one bus, the CPU can fetch only one at a time. This slowdown is called the von Neumann bottleneck.
Inside the CPU: CU, ALU and registers
- Control Unit (CU): decodes each instruction and sends control signals to the other parts.
- Arithmetic Logic Unit (ALU): does sums (add, subtract) and logic (AND, OR, compare).
- Registers: tiny, very fast stores inside the CPU.
| Register | Job |
|---|---|
| PC (Program Counter) | address of the next instruction |
| MAR (Memory Address Register) | address about to be used in memory |
| MDR (Memory Data Register) | data or instruction just read from, or about to be written to, memory |
| CIR (Current Instruction Register) | the instruction being decoded |
| ACC (Accumulator) | result of the ALU's latest calculation |
The fetch-decode-execute cycle
- Fetch: copy PC into MAR. Send the address on the address bus with a "read" signal. The instruction comes back on the data bus into MDR, then is copied to CIR. Add 1 to PC.
- Decode: the CU splits the instruction into an opcode (what to do, e.g. ADD) and an operand (what to use, e.g. address 6).
- Execute: do it. Load a value, let the ALU calculate, store a result, or jump to a new address by changing PC.
Then the cycle starts again, until a HALT instruction.
Worked trace
Memory: 0: LOAD 5, 1: ADD 6, 2: STORE 7, 3: HALT, 5: 12, 6: 30. After cycle 1, ACC = 12. After cycle 2, ACC = 42. After cycle 3, cell 7 = 42. Cycle 4 stops the program.
What makes a CPU faster?
- Clock speed: how many cycles per second, in hertz. 3 GHz = 3 billion ticks per second. Faster clock = more instructions per second, but more heat.
- Number of cores: each core is a full processor. A quad-core CPU can run 4 jobs at once, but only if the software can split the work. Two cores are not always twice as fast.
- Cache: a small, very fast memory inside the CPU that keeps recently used data and instructions. More cache = fewer slow trips to RAM. Levels: L1 (smallest, fastest), L2, L3.
Machine language and different designs
Machine code and assembly
The CPU only understands machine code: binary instructions such as 0001 0101 (opcode 0001 = LOAD, operand 0101 = 5). Assembly language writes the same instruction as LOAD 5, and an assembler turns it into binary. Each CPU family has its own instruction set.
Variation in architecture
- Harvard architecture: separate memories and buses for instructions and data, so both can be fetched at once. Used in many microcontrollers and inside CPU caches.
- CISC (Complex Instruction Set): many powerful instructions, some taking several cycles (e.g. most laptop and desktop chips).
- RISC (Reduced Instruction Set): fewer, simple instructions, usually one cycle each, low power (e.g. most phone chips).
- Multi-core and GPUs: many cores working in parallel, good for graphics and AI.
Embedded systems
An embedded system is a small computer built into a larger device to do one job: a washing machine, microwave, car brakes, a fitness band. They are cheap, small, use little power and are hard to reprogram.
Try it
Play "human CPU" with a friend. Write 8 numbered cards: 0: LOAD 5, 1: ADD 6, 2: STORE 7, 3: HALT, 5: 7, 6: 9, 7: empty. One person is the PC and points to a card; the other fetches it, reads it out (decode) and does it with a calculator (execute). What ends up on card 7? Then check in the 3D by setting cell 5 = 7 and cell 6 = 9.
Key formulas and definitions
- Stored program: instructions and data share one memory (von Neumann)
- Fetch → Decode → Execute → repeat
- Fetch: MAR ← PC; MDR ← memory[MAR]; CIR ← MDR; PC ← PC + 1
- Instruction = opcode (what to do) + operand (what to use)
- 1 GHz = 1,000,000,000 clock cycles per second
- Address bus: one way (CPU → memory); data bus: both ways
Worked examples
1. Trace the program 0: LOAD 5, 1: ADD 6, 2: STORE 7, 3: HALT with cell 5 = 12 and cell 6 = 30. Give PC and ACC after each cycle.
Cycle 1: PC = 1, ACC = 12. Cycle 2: PC = 2, ACC = 42. Cycle 3: PC = 3, ACC = 42 and cell 7 = 42. Cycle 4: PC = 4, HALT, the program stops.
2. A CPU runs at 2.5 GHz. How many clock cycles does it make in one second?
2.5 × 10⁹ = 2,500,000,000 cycles per second.
3. A 4 GHz CPU needs 2 cycles per instruction on average. Roughly how many instructions per second can one core run?
4 × 10⁹ ÷ 2 = 2 × 10⁹ instructions per second.
4. Why might a quad-core 2 GHz CPU not be twice as fast as a dual-core 2 GHz CPU for a game?
The extra cores only help if the game's work can be split into parallel tasks. If most of the work must happen one step after another, two cores sit idle.
5. An instruction is 8 bits: the first 4 bits are the opcode, the last 4 the operand. How many different opcodes and how many addresses are possible?
4 bits give 2⁴ = 16 opcodes and 2⁴ = 16 addresses (0 to 15).
6. Explain why a cache speeds up a loop that runs 1,000 times.
After the first pass, the loop's instructions and data are kept in the cache. The next 999 passes read them from the fast cache instead of slow RAM, so each fetch takes far less time.
Common mistakes
- Mixing up MAR and MDR: MAR holds an ADDRESS, MDR holds the DATA at that address.
- Thinking PC holds the current instruction. It holds the ADDRESS of the NEXT instruction; the CIR holds the current one.
- Saying the address bus carries data both ways. Only the data bus is two-way.
- Believing double the cores always means double the speed. Only if the work can be split in parallel.