Lesson 2 of 6 · Loss functions

How wrong was the prediction?

For a security ticket, predictions of 0.55 and 0.90 both pass a 0.50 threshold. But they give different amounts of support to the right answer. A loss function measures that difference so training has something to improve.

Step 1 · Define “better”

A correct/incorrect flag leaves out useful information

A threshold decision stays unchanged as a prediction moves from 0.55 to 0.90. We want a training score that notices this improvement, not just whether the prediction crosses 0.50.

For the cross-entropy loss used here, the supervised target is a class label. The loss is the negative logarithm of the probability assigned to that target. Because it is differentiable, backpropagation can calculate how each parameter affected it, and the optimizer can use those gradients to update the model.

Rank three predictions

The reviewed label is security, y=1

  • p = 0.90Strong support for the reviewed answer → small loss.
  • p = 0.55Correct side of 0.50, but uncertain → medium loss.
  • p = 0.10Confident support for the wrong answer → large loss.

Quick check

For y=1, which prediction should have the lowest loss?

Step 2 · Penalize probability on the wrong answer

Cross-entropy asks: how much probability reached the target?

For a binary classifier, let p be the predicted probability of y=1. If the reviewed label is 1, the correct-answer probability is p. If the label is 0, it is 1−p.

Binary cross-entropy is simply the negative logarithm of that correct-answer probability. The logarithm makes confident mistakes much more expensive than uncertain ones.

Worked calculation

Our neuron predicted 0.668 for a positive ticket

L = −log(pcorrect) = −log(0.668) ≈ 0.403

If it had predicted 0.10, the loss would be −log(0.10) ≈ 2.303. The confidently wrong prediction is punished much more.

Practical example · Loss workbench

Move probability toward and away from the reviewed answer

Start with label 1 and drag the prediction toward 1. Then switch the label to 0 without changing the prediction. Notice that loss cares about the reviewed target, not simply whether the raw probability is high.

Binary cross-entropy

Probability assigned to the reviewed answer:

Step 3 · Combine examples deliberately

A batch loss needs a reduction rule

This binary classifier produces one loss per example. Causal-language-model SFT usually produces one loss per unmasked target token, then reduces over the valid tokens. A mean keeps the scale roughly stable when the number of valid targets changes. A sum grows with that count.

Three-example classifier batch

mean([0.20, 0.50, 1.10]) = 1.80 / 3 = 0.60

Duplicating all three rows would double the sum to 3.60 but leave the mean at 0.60. That distinction matters when the learning rate assumes one reduction and the code uses another.

Quick check

Which reduction keeps the duplicated batch at loss 0.60?

Python bridge · classifier loss to response-token SFT
p_correct = probability if label == 1 else 1 - probability
p_correct = np.clip(p_correct, 1e-12, 1 - 1e-12)
loss = -np.log(p_correct)
batch_loss = per_example_losses.mean()

token_losses = cross_entropy(shifted_logits, target_ids, reduction="none")
sft_loss = (token_losses * response_mask).sum() / response_mask.sum()

Clipping avoids taking log(0). The SFT response mask excludes prompt and padding positions in a response-only recipe; other recipes may deliberately train additional positions.

Your turn

Explain it in your own words.

Not checked

Here p means the predicted probability that a ticket belongs in security. For a reviewed security ticket (label 1), why does p = 0.9 have less cross-entropy loss than p = 0.1? If the reviewed label is instead 0, which prediction has less loss, and why?

Answer the question in your own words. A short explanation is enough.

Draft saves on this device0/800 characters

Work through a hint