Lesson 3 of 6 · Gradients and backpropagation
Which way should a weight move?
Loss tells us how a prediction compares with its target. To improve it, we need one more piece of information: would a small increase in a weight make the loss rise or fall? That is what a derivative tells us.
Step 1 · Read the slope here
Read the slope at the current weight
Plot a weight w horizontally and loss L vertically. The derivative dL/dw is the slope at the current weight: approximately how much loss changes per unit of a small weight change.
A positive derivative means a small increase in the weight raises loss. A negative derivative means a small increase lowers loss. A derivative near zero means the curve is locally flat. It does not tell you where the global minimum is.
Worked interpretation
At w = −2, the derivative is about −0.917
This curve averages loss over six examples with z=xw and zero bias. It is a new dataset, not the single positive ticket from the loss lesson. Some examples have label 0 and others label 1, so increasing the weight indefinitely does not keep improving the average.
dL/dw ≈ −0.917
Moving w slightly upward should lower loss because the slope is negative. Gradient descent therefore subtracts the negative number, which moves the weight to the right.
After moving, evaluate the derivative at the new weight. The old slope describes the old point; the slope may be different where you landed. Checking again lets the next update use the current direction and steepness.
Predict the direction
If dL/dw is negative, which small weight change should lower loss?
Step 2 · Follow the computation backward
Backpropagation is organized chain rule
A weight affects the logit, the logit affects the probability, and the probability affects the loss. The chain rule combines these rates of change by multiplying them. Backpropagation works backward through the computation, reusing intermediate derivatives; where several paths contribute, their contributions add.
It computes gradients; it does not update parameters. The optimizer will use those gradients in the next lesson.
One-neuron chain
w → z = xw + b → p = sigmoid(z) → L
dL/dw = (dL/dz) × (dz/dw) = (p − y) × x
Where does p − y come from?
For binary cross-entropy, dL/dp = (p−y) / [p(1−p)]. For sigmoid, dp/dz = p(1−p). Multiplying these derivatives cancels the p(1−p) factors, leaving dL/dz = p−y. Finally, dz/dw=x, giving dL/dw=(p−y)x.
For p=0.668, y=1, and x₁=1, the gradient is about −0.332. For x₂=0, the gradient flowing into w₂ is zero for this example.
Quick check
For this example, what gradient reaches a weight whose input feature is zero?
Step 3 · Test the derivative
Check a derivative by nudging the weight
Compute loss with the weight nudged slightly up and slightly down. Their difference, divided by the distance between the two weights, estimates the derivative without backpropagation.
Centered finite difference
dL/dw ≈ [L(w + ε) − L(w − ε)] / (2ε)
In the included Python example, the analytical gradient −0.331812227832 and finite-difference estimate −0.331812227805 agree to many decimal places.
Practical example · Loss-curve explorer
Predict the sign, then take one local step
Move the weight and predict whether its slope is positive or negative. Then take a step and compare the new loss. The learning rate scales the move; the cyan ring marks the next point. This explorer clips steps to the plotted range −5 to 5. The next lesson shows unclipped updates.
analytical = (probability - label) * feature
numerical = (loss(w + eps) - loss(w - eps)) / (2 * eps)
assert np.allclose(analytical, numerical, atol=1e-6)The assertion checks the derivative. The optimizer—not backpropagation—will decide how far to move the weight.
Your turn
Explain it in your own words.
At weight w = −2, the loss derivative is −0.917. To lower loss locally, should we move the weight a little higher or lower? Why evaluate the derivative at the new weight before choosing another move?
Answer the question in your own words. A short explanation is enough.
Your feedback
Work through a hint
Hint 1
A negative slope slopes down as you move right. Which direction increases w?
Hint 2
The value −0.917 was calculated at −2. Does it describe every other point on the curve?
A worked explanation
Increase w a little to move downhill locally. Then calculate the derivative where you landed, because the slope there may differ from the old one.