Optional reference after lesson 6 · Attention in detail
Attention weights and context costs
This optional reference opens up the attention calculation used by the Shakespeare model. Complete the six project lessons first; you do not need these formulas to train it.
You have seen how earlier context can change a prediction. Now inspect how attention computes a weighted mixture, why longer input costs more, and what cached computations can be reused while generating.
Step 1 · Build context from the current input
Attention combines information from several positions
Each position’s representation is projected into three vectors: a query, a key, and a value. To predict the next token, the final input position’s query is compared with keys at permitted positions, including itself. These comparisons produce scores; the values contain the information to combine.
Softmax turns the query-key scores into nonnegative weights that sum to one. The attention output is the weighted sum of value vectors. Different heads can learn different mixtures, and later layers can recombine them.
One illustrative head
scores = query · keys / √d → softmax → attention weights
context vector = Σ weightᵢ × valueᵢ
Here d is the number of features in one head, and Σ means to add the weighted vectors together.
A routing-oriented head might place its largest weight on “credentials,” while still mixing information from “route,” “to,” and other earlier positions. The weights are computation, not a guaranteed human explanation of model reasoning.
Quick check
What does a standard attention head return?
Practical example · Context-mechanism comparison
Compare a fixed window with two attention patterns
With three visible previous tokens, “credentials” is outside the fixed window. Switch to attention and compare the two prepared weight patterns. They illustrate how different mixtures use earlier information; these weights are not calculated by a live transformer.
Step 2 · Preserve order and prevent leakage
A decoder may use earlier tokens, never future targets
A causal mask prevents a prediction position from attending to tokens that come after it. Without the mask, training could leak the answer from the future.
Attention by itself does not know token order. Positional information—learned, sinusoidal, rotary, or another scheme—lets the model distinguish “dog bites person” from “person bites dog.”
Causal visibility
- Predict token 3May use tokens 1–2, not token 3 or anything after it.
- Predict token 8May use tokens 1–7, subject to the model’s context limit.
- TrainingComputes many positions in parallel while the mask enforces the same rule.
Quick check
What prevents next-token training from reading the answer to its right?
Step 3 · Connect attention to the rest of the model
Attention is one component of a transformer block
A transformer repeats blocks of attention and a small feed-forward network (MLP). Attention mixes information across positions; the MLP transforms each position’s vector. Residual connections add a block’s input back to its output, and normalization controls the scale of intermediate values. The exact ordering varies by architecture. A final output layer produces one score per vocabulary token.
From tokens to the training loss
token IDs → embeddings + position → [attention → residual/norm → MLP → residual/norm] × N → vocabulary logits
vocabulary logits + observed next-token target → cross-entropy loss
As in the Shakespeare project, the observed next token supplies the target. Training adjusts the embeddings and all the blocks together.
Serving check
What does the KV cache primarily avoid recomputing during decode?
scores = (queries @ keys.T) / sqrt(head_dimension)
scores = scores.masked_fill(future_positions, -inf)
weights = stable_softmax(scores)
context = weights @ valuesReal implementations batch sequences and heads, use optimized kernels, and carefully manage cache layout. The four operations remain visible underneath.