Lesson 2 of 6 · Loss functions
How wrong was the prediction?
For a security ticket, predictions of 0.55 and 0.90 both pass a 0.50 threshold. But they give different amounts of support to the right answer. A loss function measures that difference so training has something to improve.
Step 1 · Define “better”
A correct/incorrect flag leaves out useful information
A threshold decision stays unchanged as a prediction moves from 0.55 to 0.90. We want a training score that notices this improvement, not just whether the prediction crosses 0.50.
For the cross-entropy loss used here, the supervised target is a class label. The loss is the negative logarithm of the probability assigned to that target. Because it is differentiable, backpropagation can calculate how each parameter affected it, and the optimizer can use those gradients to update the model.
Rank three predictions
The reviewed label is security, y=1
- p = 0.90Strong support for the reviewed answer → small loss.
- p = 0.55Correct side of 0.50, but uncertain → medium loss.
- p = 0.10Confident support for the wrong answer → large loss.
Quick check
For y=1, which prediction should have the lowest loss?
Step 2 · Penalize probability on the wrong answer
Cross-entropy asks: how much probability reached the target?
For a binary classifier, let p be the predicted probability of y=1. If the reviewed label is 1, the correct-answer probability is p. If the label is 0, it is 1−p.
Binary cross-entropy is simply the negative logarithm of that correct-answer probability. The logarithm makes confident mistakes much more expensive than uncertain ones.
Worked calculation
Our neuron predicted 0.668 for a positive ticket
L = −log(pcorrect) = −log(0.668) ≈ 0.403
If it had predicted 0.10, the loss would be −log(0.10) ≈ 2.303. The confidently wrong prediction is punished much more.
Practical example · Loss workbench
Move probability toward and away from the reviewed answer
Start with label 1 and drag the prediction toward 1. Then switch the label to 0 without changing the prediction. Notice that loss cares about the reviewed target, not simply whether the raw probability is high.
Binary cross-entropy
Probability assigned to the reviewed answer:
Step 3 · Combine examples deliberately
A batch loss needs a reduction rule
This binary classifier produces one loss per example. Causal-language-model SFT usually produces one loss per unmasked target token, then reduces over the valid tokens. A mean keeps the scale roughly stable when the number of valid targets changes. A sum grows with that count.
Three-example classifier batch
mean([0.20, 0.50, 1.10]) = 1.80 / 3 = 0.60
Duplicating all three rows would double the sum to 3.60 but leave the mean at 0.60. That distinction matters when the learning rate assumes one reduction and the code uses another.
Quick check
Which reduction keeps the duplicated batch at loss 0.60?
p_correct = probability if label == 1 else 1 - probability
p_correct = np.clip(p_correct, 1e-12, 1 - 1e-12)
loss = -np.log(p_correct)
batch_loss = per_example_losses.mean()
token_losses = cross_entropy(shifted_logits, target_ids, reduction="none")
sft_loss = (token_losses * response_mask).sum() / response_mask.sum()
Clipping avoids taking log(0). The SFT response mask excludes prompt and padding positions in a response-only recipe; other recipes may deliberately train additional positions.
Your turn
Explain it in your own words.
Here p means the predicted probability that a ticket belongs in security. For a reviewed security ticket (label 1), why does p = 0.9 have less cross-entropy loss than p = 0.1? If the reviewed label is instead 0, which prediction has less loss, and why?
Answer the question in your own words. A short explanation is enough.
Your feedback
Work through a hint
Hint 1
Here p is always the probability of security, not necessarily the probability of the correct answer.
Hint 2
If p is 0.1, how much probability is left for not-security?
A worked explanation
For label 1, p=0.9 gives the correct class 0.9. For label 0, p=0.1 leaves 0.9 for the correct negative class. More probability on the target means less cross-entropy loss.