Lesson 4 of 6 · The optimizer loop
Take a step, then calculate again
You have a gradient. Now use it to change a weight, calculate the new loss, and repeat. The optimizer supplies the update rule; its learning rate determines how large each move will be.
Step 1 · Subtract the gradient
Multiply the gradient by the learning rate
Plain stochastic gradient descent uses one compact rule. The gradient supplies direction and relative steepness. The learning rate η scales that information into a step.
One SGD update
wnext = w − η(dL/dw)
wnext = −2 − 2(−0.917) ≈ −0.166
The negative gradient makes the subtraction add. A smaller learning rate would move less; an excessively large one can jump beyond the useful region.
Quick check
Which update is gradient descent?
Step 2 · Repeat the calculation
A training step has a specific order
For each batch, training runs a forward pass, calculates loss, backpropagates gradients, applies an optimizer step, and clears or replaces accumulated gradients. Then it repeats with another batch.
A batch is one group of examples. A step is one optimizer update. An epoch is one pass through the training dataset. Framework method optimizer.step() means “apply the update”; it is not a mathematical step function.
Training-loop order
batch → forward → loss → backward → optimizer step → clear gradients
Evaluation and inference normally stop after the forward pass. They do not call backward or mutate deployed weights.
Quick check
What normally happens immediately after backpropagation?
Practical example · Optimizer trace
Run the same gradients with different learning rates
This uses the previous lesson’s six-example average loss with z=xw. Reset at w=−2 before each run. Compare six steps at 0.2, 2.0, and 40.0: how much does loss fall, and does any step raise it? These rates illustrate this curve, not recommended settings for an LLM.
| Step | w before | gradient | w after | loss after |
|---|
Step 3 · Checkpoint the process
Why saving weights is not enough to resume Adam
Adam keeps running averages of gradients and squared gradients to adjust its updates. A learning-rate scheduler changes the rate as training progresses. If you restore only the weights, those averages and the scheduler’s position are missing: the next update can differ even on the same batch.
A resumable training checkpoint
- Model stateBase weights plus any trainable adapter weights.
- Optimizer stateMomentum or Adam running averages.
- Schedule stateCurrent learning rate and global step.
- Run stateDataset position and random-number state when exact resumption matters.
optimizer.zero_grad()
prediction = model(batch_inputs)
loss = loss_fn(prediction, batch_labels)
loss.backward()
optimizer.step()In this PyTorch-style loop, zero_grad() clears old gradients before the next backward pass. backward() calculates new gradients; step() uses them to update the parameters.
Your turn
Explain it in your own words.
A backward pass computes dL/dw = −0.5. Plain SGD uses w_next = w − learning_rate × dL/dw. What information did backward produce, what does the optimizer change, and how would doubling the learning rate change this one update?
Answer the question in your own words. A short explanation is enough.
Your feedback
Work through a hint
Hint 1
Look at the formula: which value is the derivative, and which value is being changed?
Hint 2
Keep −0.5 fixed and double only the learning rate. What happens to their product?
A worked explanation
Backward calculated the loss slope; it did not move the weight. SGD adds 0.5 times the learning rate to w here. Doubling that rate doubles this one increase.