Lesson 4 of 6 · The optimizer loop

Take a step, then calculate again

You have a gradient. Now use it to change a weight, calculate the new loss, and repeat. The optimizer supplies the update rule; its learning rate determines how large each move will be.

Step 1 · Subtract the gradient

Multiply the gradient by the learning rate

Plain stochastic gradient descent uses one compact rule. The gradient supplies direction and relative steepness. The learning rate η scales that information into a step.

One SGD update

wnext = w − η(dL/dw)

wnext = −2 − 2(−0.917) ≈ −0.166

The negative gradient makes the subtraction add. A smaller learning rate would move less; an excessively large one can jump beyond the useful region.

Quick check

Which update is gradient descent?

Step 2 · Repeat the calculation

A training step has a specific order

For each batch, training runs a forward pass, calculates loss, backpropagates gradients, applies an optimizer step, and clears or replaces accumulated gradients. Then it repeats with another batch.

A batch is one group of examples. A step is one optimizer update. An epoch is one pass through the training dataset. Framework method optimizer.step() means “apply the update”; it is not a mathematical step function.

Training-loop order

batch → forward → loss → backward → optimizer step → clear gradients

Evaluation and inference normally stop after the forward pass. They do not call backward or mutate deployed weights.

Quick check

What normally happens immediately after backpropagation?

Practical example · Optimizer trace

Run the same gradients with different learning rates

This uses the previous lesson’s six-example average loss with z=xw. Reset at w=−2 before each run. Compare six steps at 0.2, 2.0, and 40.0: how much does loss fall, and does any step raise it? These rates illustrate this curve, not recommended settings for an LLM.

Step Weight Loss Current gradient
Stepw beforegradientw afterloss after

Step 3 · Checkpoint the process

Why saving weights is not enough to resume Adam

Adam keeps running averages of gradients and squared gradients to adjust its updates. A learning-rate scheduler changes the rate as training progresses. If you restore only the weights, those averages and the scheduler’s position are missing: the next update can differ even on the same batch.

A resumable training checkpoint

  • Model stateBase weights plus any trainable adapter weights.
  • Optimizer stateMomentum or Adam running averages.
  • Schedule stateCurrent learning rate and global step.
  • Run stateDataset position and random-number state when exact resumption matters.
Python bridge · explicit training-loop responsibilities
optimizer.zero_grad()
prediction = model(batch_inputs)
loss = loss_fn(prediction, batch_labels)
loss.backward()
optimizer.step()

In this PyTorch-style loop, zero_grad() clears old gradients before the next backward pass. backward() calculates new gradients; step() uses them to update the parameters.

Your turn

Explain it in your own words.

Not checked

A backward pass computes dL/dw = −0.5. Plain SGD uses w_next = w − learning_rate × dL/dw. What information did backward produce, what does the optimizer change, and how would doubling the learning rate change this one update?

Answer the question in your own words. A short explanation is enough.

Draft saves on this device0/800 characters

Work through a hint