From brain cells to artificial neurons
Your brain has billions of nerve cells. Each one receives signals, and if the total signal is strong enough, it sends a signal on. Computer scientists copied this idea (not the biology) to make an artificial neuron.
- Inputs (x): numbers that come in, for example pixel brightness.
- Weights (w): how important each input is. A big weight means "listen to this input a lot". A negative weight means "this input argues against".
- Bias (b): an extra number that shifts how easy it is to fire.
- Weighted sum: z = w₁x₁ + w₂x₂ + … + b.
- Activation function: turns z into the output.
The very first model of this kind, the perceptron (1958), used a simple yes/no rule: output 1 if z ≥ 0, otherwise 0.
Activation functions: when does a neuron fire?
Without an activation function, a network could only draw straight lines. Activation adds a bend, so the network can learn curved, complex patterns.
- Step: 0 or 1. Simple but hard to train.
- Sigmoid: squashes any number into 0 to 1. σ(z) = 1 ÷ (1 + e−z). σ(0) = 0.5.
- ReLU: output = z if z > 0, else 0. Fast and very popular in deep networks.
- Softmax: used in the output layer to turn scores into probabilities that add up to 1.
Layers: input, hidden and output
Neurons are arranged in layers. The input layer just holds the data (for a 28 × 28 picture, 784 numbers). The hidden layers find patterns: early ones notice simple things like edges, later ones combine them into shapes. The output layer gives the answer, for example 10 neurons for the digits 0–9.
If every neuron in one layer connects to every neuron in the next, the layer is fully connected (dense). The number of weights between a layer of 3 and a layer of 4 is 3 × 4 = 12, plus 4 biases.
How a network learns: loss, backpropagation and gradient descent
- Forward pass: put an example in, let the numbers flow to the output.
- Loss: compare the output with the correct label. The loss is a number that is big when the guess is bad.
- Backpropagation: work backwards from the output to find how much each weight caused the error.
- Gradient descent: change each weight a small step in the direction that lowers the loss. The step size is the learning rate.
- Repeat for all examples many times. One full pass over the training data is an epoch.
Too few examples or too many epochs can cause overfitting: the network memorises the training data and does badly on new data. That is why we keep a separate test set.
Deep learning: CNN, RNN and frameworks
Deep learning means neural networks with many hidden layers, trained on large data with fast chips (GPUs).
- CNN (convolutional neural network): a small grid of weights called a filter (kernel) slides across the image. At each place it multiplies and adds, making a feature map that lights up where an edge or shape is found. Used for face unlock, medical scans, self-driving cars.
- RNN (recurrent neural network): has a loop, so the output for one word is fed back in with the next word. It has a short memory. Used for text, speech, music. Newer transformer models use attention instead of loops and power today's chatbots and translators.
- Frameworks: ready-made code libraries such as TensorFlow, Keras and PyTorch. You describe the layers; the framework does the maths and backpropagation for you.
Try it: be a neuron with paper and pencil
Decide whether to go out to play. Inputs: x₁ = 1 if it is sunny, x₂ = 1 if homework is done. Weights: w₁ = 2, w₂ = 3. Bias b = −4. Rule: fire (go out) if z ≥ 0.
- Sunny, homework not done: z = 2 + 0 − 4 = −2 → stay in.
- Sunny, homework done: z = 2 + 3 − 4 = 1 → go out!
Now change the bias to −1 and test again. Then open free play in the 3D above and do the same with sliders.
Key formulas and definitions
- Weighted sum: z = w₁x₁ + w₂x₂ + … + wₙxₙ + b
- Output: a = f(z), f = activation function
- Sigmoid: σ(z) = 1 ÷ (1 + e^(−z)); ReLU: f(z) = max(0, z)
- Weights between layers of m and n neurons = m × n (+ n biases)
- Gradient descent: new w = old w − learning rate × (slope of loss for w)
Worked examples
1. A neuron has inputs x₁ = 2, x₂ = 1, weights w₁ = 0.5, w₂ = −1 and bias b = 0.5. Find z and the ReLU output.
z = 0.5×2 + (−1)×1 + 0.5 = 1 − 1 + 0.5 = 0.5. ReLU(0.5) = 0.5.
2. A step neuron fires if z ≥ 0. Inputs (1, 1), weights (1, 1), bias −1.5. Does it fire? What does this neuron compute?
z = 1 + 1 − 1.5 = 0.5 ≥ 0 → fires. With (1,0) or (0,1): z = −0.5 → no; (0,0): z = −1.5 → no. It fires only when both inputs are 1: it acts as an AND gate.
3. How many weights and biases does a fully connected network 4 → 5 → 3 have?
Weights: 4×5 + 5×3 = 20 + 15 = 35. Biases: one per non-input neuron = 5 + 3 = 8. Total parameters = 43.
4. A weight is 0.8. The slope of the loss for this weight is 2 and the learning rate is 0.1. Find the new weight.
new w = 0.8 − 0.1 × 2 = 0.6. The weight moves down because increasing it would raise the loss.
5. A 3×3 filter slides one step at a time over a 6×6 image with no padding. What size is the feature map?
Positions in each direction = 6 − 3 + 1 = 4. Feature map = 4 × 4 (as in the 3D, step 5).
6. Find the sigmoid output when z = 0 and when z is very large.
σ(0) = 1 ÷ (1 + 1) = 0.5. For very large z, e^(−z) ≈ 0, so σ ≈ 1. Sigmoid always stays between 0 and 1.
Common mistakes
- Thinking an artificial neuron is a real brain cell. It is only a maths idea inspired by the brain.
- Forgetting to add the bias when finding z.
- Counting the input layer as a layer of neurons with weights. It only holds the data; weights sit on the connections after it.
- Believing more training is always better. Too much training on little data leads to overfitting.