Lesson 3 of 6 · Gradients and backpropagation

Which way should a weight move?

Loss tells us how a prediction compares with its target. To improve it, we need one more piece of information: would a small increase in a weight make the loss rise or fall? That is what a derivative tells us.

Step 1 · Read the slope here

Read the slope at the current weight

Plot a weight w horizontally and loss L vertically. The derivative dL/dw is the slope at the current weight: approximately how much loss changes per unit of a small weight change.

A positive derivative means a small increase in the weight raises loss. A negative derivative means a small increase lowers loss. A derivative near zero means the curve is locally flat. It does not tell you where the global minimum is.

Worked interpretation

At w = −2, the derivative is about −0.917

This curve averages loss over six examples with z=xw and zero bias. It is a new dataset, not the single positive ticket from the loss lesson. Some examples have label 0 and others label 1, so increasing the weight indefinitely does not keep improving the average.

dL/dw ≈ −0.917

Moving w slightly upward should lower loss because the slope is negative. Gradient descent therefore subtracts the negative number, which moves the weight to the right.

After moving, evaluate the derivative at the new weight. The old slope describes the old point; the slope may be different where you landed. Checking again lets the next update use the current direction and steepness.

Predict the direction

If dL/dw is negative, which small weight change should lower loss?

Step 2 · Follow the computation backward

Backpropagation is organized chain rule

A weight affects the logit, the logit affects the probability, and the probability affects the loss. The chain rule combines these rates of change by multiplying them. Backpropagation works backward through the computation, reusing intermediate derivatives; where several paths contribute, their contributions add.

It computes gradients; it does not update parameters. The optimizer will use those gradients in the next lesson.

One-neuron chain

w → z = xw + b → p = sigmoid(z) → L

dL/dw = (dL/dz) × (dz/dw) = (p − y) × x

Where does p − y come from?

For binary cross-entropy, dL/dp = (p−y) / [p(1−p)]. For sigmoid, dp/dz = p(1−p). Multiplying these derivatives cancels the p(1−p) factors, leaving dL/dz = p−y. Finally, dz/dw=x, giving dL/dw=(p−y)x.

For p=0.668, y=1, and x₁=1, the gradient is about −0.332. For x₂=0, the gradient flowing into w₂ is zero for this example.

Quick check

For this example, what gradient reaches a weight whose input feature is zero?

Step 3 · Test the derivative

Check a derivative by nudging the weight

Compute loss with the weight nudged slightly up and slightly down. Their difference, divided by the distance between the two weights, estimates the derivative without backpropagation.

Centered finite difference

dL/dw ≈ [L(w + ε) − L(w − ε)] / (2ε)

In the included Python example, the analytical gradient −0.331812227832 and finite-difference estimate −0.331812227805 agree to many decimal places.

Practical example · Loss-curve explorer

Predict the sign, then take one local step

Move the weight and predict whether its slope is positive or negative. Then take a step and compare the new loss. The learning rate scales the move; the cyan ring marks the next point. This explorer clips steps to the plotted range −5 to 5. The next lesson shows unclipped updates.

Loss Derivative Lowest plotted loss Next weight
Python bridge · calculate, check, then use the gradient
analytical = (probability - label) * feature
numerical = (loss(w + eps) - loss(w - eps)) / (2 * eps)
assert np.allclose(analytical, numerical, atol=1e-6)

The assertion checks the derivative. The optimizer—not backpropagation—will decide how far to move the weight.

Your turn

Explain it in your own words.

Not checked

At weight w = −2, the loss derivative is −0.917. To lower loss locally, should we move the weight a little higher or lower? Why evaluate the derivative at the new weight before choosing another move?

Answer the question in your own words. A short explanation is enough.

Draft saves on this device0/800 characters

Work through a hint