Lesson 6 of 6 · Multiclass prediction

How do scores become probabilities?

Our router now chooses among security, billing, and product. It produces one score per class. Softmax converts those scores into probabilities that sum to one—the same calculation an LLM uses across possible next tokens.

Step 1 · Produce competing scores

A score of 2 is not a probability of 200%

A three-way ticket router produces scores for security, billing, and product. Logits can be negative, do not need to sum to one, and are meaningful mainly through their differences.

An LLM does the same job at a larger scale: its output head produces one logit for every token in the vocabulary at the current sequence position.

One ticket, three logits

z = [2.0, 0.5, −1.0]

Security has the largest raw score, but 2.0 is not “200% confidence.” We need a normalization that preserves ordering and returns positive values summing to one.

Quick check

Which statement about logits is true?

Step 2 · Normalize without overflow

Softmax exponentiates differences between logits

Exponentiation makes larger logits receive more probability, then division by the total makes the values sum to one. Adding the same constant to every logit does not change their differences, so it should not change the distribution.

Stable softmax subtracts the largest logit before exponentiating. That makes the largest exponent exactly e⁰=1 and avoids enormous intermediate values.

Why is a shared shift harmless? Adding c to every logit multiplies every exponential by eᶜ. That same factor appears in the numerator and denominator, so it cancels.

Stable calculation

[2.0, 0.5, −1.0] − max = [0.0, −1.5, −3.0]

softmax ≈ [0.786, 0.175, 0.039]

The result preserves the ranking and sums to one.

Numerical-stability check

What happens if 100 is added to every logit?

Practical example · Stable-softmax workbench

Change evidence, then shift every score together

First change one class logit and watch probability move between classes. Then add 100 to all three scores; the displayed logits change while the probabilities stay fixed.

Top class

Shifted logits:

Probability sum:

ClassLogitProbability

Step 3 · Score the reviewed target

Multiclass cross-entropy reads one probability

The label is the index of the reviewed class. Cross-entropy selects that class’s softmax probability and takes its negative log. The other classes still matter because softmax makes all classes compete for the same probability mass.

Security is the target

L = −log P(target=security) = −log(0.786) ≈ 0.241

During next-token training, the target class is the observed next token. The model is trained against that token even if another token was its current top prediction.

Quick check

Which probability does cross-entropy place inside −log(·)?

Python bridge · stable softmax and target loss
shifted = logits - logits.max(axis=-1, keepdims=True)
probabilities = np.exp(shifted) / np.exp(shifted).sum(axis=-1, keepdims=True)
loss = -np.log(probabilities[row, target_class])

For a batch, the final loss is normally the mean of one target loss per row.

Your turn

Explain it in your own words.

Not checked

A ticket router scores security, billing, and product as [2.0, 0.5, −1.0]. How does softmax turn these scores into probabilities? If the reviewed label is billing, which probability does cross-entropy use—even though security has the largest score?

Answer the question in your own words. A short explanation is enough.

Draft saves on this device0/800 characters

Work through a hint