Lesson 6 of 6 · Multiclass prediction
How do scores become probabilities?
Our router now chooses among security, billing, and product. It produces one score per class. Softmax converts those scores into probabilities that sum to one—the same calculation an LLM uses across possible next tokens.
Step 1 · Produce competing scores
A score of 2 is not a probability of 200%
A three-way ticket router produces scores for security, billing, and product. Logits can be negative, do not need to sum to one, and are meaningful mainly through their differences.
An LLM does the same job at a larger scale: its output head produces one logit for every token in the vocabulary at the current sequence position.
One ticket, three logits
z = [2.0, 0.5, −1.0]
Security has the largest raw score, but 2.0 is not “200% confidence.” We need a normalization that preserves ordering and returns positive values summing to one.
Quick check
Which statement about logits is true?
Step 2 · Normalize without overflow
Softmax exponentiates differences between logits
Exponentiation makes larger logits receive more probability, then division by the total makes the values sum to one. Adding the same constant to every logit does not change their differences, so it should not change the distribution.
Stable softmax subtracts the largest logit before exponentiating. That makes the largest exponent exactly e⁰=1 and avoids enormous intermediate values.
Why is a shared shift harmless? Adding c to every logit multiplies every exponential by eᶜ. That same factor appears in the numerator and denominator, so it cancels.
Stable calculation
[2.0, 0.5, −1.0] − max = [0.0, −1.5, −3.0]
softmax ≈ [0.786, 0.175, 0.039]
The result preserves the ranking and sums to one.
Numerical-stability check
What happens if 100 is added to every logit?
Practical example · Stable-softmax workbench
Change evidence, then shift every score together
First change one class logit and watch probability move between classes. Then add 100 to all three scores; the displayed logits change while the probabilities stay fixed.
Top class
Shifted logits:
Probability sum:
| Class | Logit | Probability |
|---|
Step 3 · Score the reviewed target
Multiclass cross-entropy reads one probability
The label is the index of the reviewed class. Cross-entropy selects that class’s softmax probability and takes its negative log. The other classes still matter because softmax makes all classes compete for the same probability mass.
Security is the target
L = −log P(target=security) = −log(0.786) ≈ 0.241
During next-token training, the target class is the observed next token. The model is trained against that token even if another token was its current top prediction.
Quick check
Which probability does cross-entropy place inside −log(·)?
shifted = logits - logits.max(axis=-1, keepdims=True)
probabilities = np.exp(shifted) / np.exp(shifted).sum(axis=-1, keepdims=True)
loss = -np.log(probabilities[row, target_class])For a batch, the final loss is normally the mean of one target loss per row.
Your turn
Explain it in your own words.
A ticket router scores security, billing, and product as [2.0, 0.5, −1.0]. How does softmax turn these scores into probabilities? If the reviewed label is billing, which probability does cross-entropy use—even though security has the largest score?
Answer the question in your own words. A short explanation is enough.
Your feedback
Work through a hint
Hint 1
Raw scores can be negative and need not sum to one. What does exponentiating, then dividing by a total, accomplish?
Hint 2
The model favors security, but the reviewed label is billing. Which one supplies the target for learning?
A worked explanation
Exponentiate each score and divide by the sum to obtain probabilities. The loss uses billing's probability because the target comes from the label, not the largest model score.