Backpropagation is the algorithm that teaches a neural network. It answers one question: “Which weights caused the error, and by how much?” It works in two parts — a forward pass (make a prediction) and a backward pass (trace the error backward through the network, using the chain rule, to give every weight its own “blame score” or gradient). The optimizer then nudges each weight in the direction that reduces the error. This loop — predict, measure, blame, adjust — repeated millions of times is how every modern AI, from GPT to image classifiers, learns.
1. The Big Picture: How a Neural Network Learns
Training a neural network is a repeating loop of 4 steps:
Forward pass — feed input data through the network to get a prediction.
Compute loss — measure how wrong the prediction was.
Backward pass (backpropagation) — figure out how much each weight contributed to that error.
Update weights — nudge every weight in the direction that reduces the error.
flowchart LR
I["Input"] --> FP["Forward Pass"] --> P["Prediction (y-hat)"] --> C{"Compare with truth"} --> L["Loss"]
L --> BP["Backward Pass"] --> G["Gradients (blame scores)"] --> U["Update Weights"] --> I
Backpropagation is only step 3
Backpropagation computes the gradients (blame scores). Gradient descent / optimizers (step 4) use those gradients to update the weights. People often mix these up — they are two separate pieces.
Analogy: Blindfolded Darts
Imagine throwing darts blindfolded at a target.
The forward pass is your throw (a guess).
The loss is how far off you were.
Backpropagation is the precise feedback: “your elbow was too high, your wrist turned too far left” — telling each muscle exactly how it contributed.
The weight update is adjusting your throw accordingly.
A real neural network has millions or billions of weights. After a wrong prediction, the network needs to know, for every single weight:
Did this weight cause the error?
By how much?
Which direction should it change to reduce the error?
Computing this naively (one calculation per weight, from scratch) is impossibly expensive. With 1 million weights, that’s 1 million separate full runs. Backpropagation computes all of them in roughly the cost of 2 runs — one forward, one backward.
3. The Forward Pass (Quick Recap)
Before backpropagation can run, the network must make a prediction. Data flows layer by layer: each neuron multiplies its inputs by its weights, adds a bias, passes the result through an activation function (like ReLU), and sends it to the next layer.
The forward pass also saves intermediate values (activations) at every layer — the backward pass needs them later to compute gradients. This is why training uses much more memory than inference.
4. The Loss: How Wrong Are We?
The loss function measures the difference between the prediction y^ and the true answer y. A common one is Mean Squared Error (MSE): L=n1∑(y−y^)2
Training is just minimizing this loss. Visualize it as a landscape — for one weight, it’s a parabola: the bottom is the best value for that weight.
5. The Secret Ingredient: The Chain Rule
The chain rule is the math behind backpropagation. It says: if a value depends on another value, which depends on another, you can compute the sensitivity by multiplying the local sensitivities along the path.
flowchart LR
W["Weight w"] -->|"(dz/dw) — the input flowing through this weight"| Z["z"]
Z -->|"(dŷ/dz) — slope of the activation function"| Y["Prediction ŷ"]
Y -->|"(dLoss/dŷ) — how wrong the model was"| L["Loss"]
L ==>|"dLoss/dw = dLoss/dŷ · dŷ/dz · dz/dw"| W
A neural network is exactly this: a long chain of nested functions. The chain rule lets us walk backward from the loss, one layer at a time, reusing already-computed values. See 1. Artificial Neural Networks (ANNs) for the full derivation.
6. The Backward Pass: Blame, Layer by Layer
Backpropagation starts at the loss and walks backward to the input. At each node it follows one recipe:
Upstream gradient × local derivative = gradient to pass down
Worked Example with Real Numbers
Take a tiny network: one input x = 1, one hidden unit, weight w1 = 0.5, output weight w2 = -2, target t = 1, ReLU activation.
Forward pass:a = 0.5, h = ReLU(0.5) = 0.5, o = (-2)(0.5) = -1, loss L = (-1 - 1)² = 4
Backward pass (start at the loss, walk back):
flowchart LR
L["Loss L = 4"] -->|"dL/do = 2(o - t) = -4"| O["Output o = -1"]
O -->|"dL/dw2 = dL/do · h = -2 — blame for output weight"| W2["w2"]
O -->|"dL/dh = dL/do · w2 = +8 — pass blame down"| H["Hidden h = 0.5"]
H -->|"dL/da = dL/dh · relu' = +8 — ReLU slope is 1"| A["Activation a = 0.5"]
A -->|"dL/dw1 = dL/da · x = +8 — blame for input weight"| W1["w1"]
What do these numbers mean?
dL/dw1 = +8 — increasing w1 increases the loss, so gradient descent will decrease it.
dL/dw2 = -2 — increasing w2 decreases the loss, so we increase it.
The sign tells the direction; the size tells how sensitive the loss is to that weight.
This is the same pattern as the chain rule equation in section 5 — just applied systematically, layer by layer, from output to input.
7. The Update: Gradient Descent
Once we have every gradient, the optimizer moves each weight a small step downhill:
w←w−η⋅∂w∂L
η (eta) is the learning rate — how big each step is.
Too big → you overshoot and bounce around. Too small → learning takes forever.
flowchart LR
W0["w = 10"] -->|"step: w - η · dL/dw"| W1["w = 7.5"]
W1 -->|"step"| W2["w = 5.8"]
W2 -->|"step"| W3["w = 5.1"]
W3 -->|"step"| W4["w = 5.0 (minimum)"]
The backward pass reuses intermediate values saved during the forward pass — no gradient is ever computed from scratch twice.
One forward + one backward pass gives gradients for every weight, no matter how many there are.
The backward pass costs roughly 2× the forward pass, so a full training step is about 3× one inference. This is why training billion-parameter models is even possible.
In practice you never do this by hand
Frameworks like PyTorch and TensorFlow build a computation graph during the forward pass and apply backpropagation automatically with one line of code:
loss.backward() # computes every gradient via autogradoptimizer.step() # uses them to update weights
9. When Things Go Wrong
Vanishing gradients — the blame signal shrinks as it travels back through many layers (the gradient is a product of many small factors). Early layers stop learning. Fixes: ReLU activations, residual connections (like in Transformers).
Exploding gradients — the signal grows out of control. Fixes: lower learning rate, gradient clipping.
Dead ReLU — if a neuron’s input is negative, ReLU’s slope is 0, so no gradient flows through it — that neuron stops learning entirely.