Lesson 2 of 6 · From counts to transformers
How transformers use context to predict new text
Our trigram sees only the last two tokens, so it misses earlier clues. It also has no estimate when that exact two-token context never appeared in its training text. We want a model that can use more of the text and make useful guesses for new combinations of familiar tokens.
Why not just count a longer context?
In lesson 1, we counted what followed the last two tokens. Try that method on these four new training sentences. In this deliberately small example, crashes are followed by “urgent” and typos by “minor.”
the crash in the new build is urgentthe typo in the new build is minorthe crash in the latest release is urgentthe typo in the latest release is minor
Now complete: the crash in the new release is …
You have evidence for “urgent,” even though this exact prefix never occurred. All its words appeared in the training text; their combination is new.
Give the table more context
Change “crash” to “typo,” then widen the window. Can the table use that difference?
Exact context: release is
urgent: 1/2 (50%); minor: 1/2 (50%)
The last two tokens are identical for crash and typo. The table must give both the same probabilities.
These are exact counts from the four displayed lines, calculated in Python. Words stand in for tokens; we inspect a prediction inside a sequence. Each training line also has beginning/end markers, as in lesson 1.
See all three window sizes together
| Count-model input | Exact context | Result from the four lines |
|---|---|---|
| Last 2 tokens | release is | minor: 1/2 (50%); urgent: 1/2 (50%) |
| Last 3 tokens | new release is | No matches → no estimate |
| Whole prefix | the crash in the new release is | No matches → no estimate |
A short window loses the clue; a long exact match loses the evidence. “Release is” has observations but excludes “crash.” The whole prefix includes “crash” but has no matching observations. More text can help, yet new phrasings keep creating new table entries to fill.
You could write a rule that sees “crash” and predicts “urgent.” But try “the crash was fixed, so the ticket is …”: “closed” may now fit better. We want the model to learn which relationships matter, rather than hand-write a rule for each situation.
Keep the task. Replace the lookup table.
Our count model takes a context and returns probabilities for the next token. A neural network can do the same job: give it a context, and ask how likely each next token is. It is a function with adjustable numbers. The same text that taught our table what to count can teach the network how to predict.
For this comparison, both models see just the last two tokens.
Count model
When predicting
build isNeural model
When predicting
build isDuring training · the same observed example
In the crash in the new build is urgent, the context build is was followed by urgent.
Add 1 to the count for build is → urgent.
Compare its prediction with urgent. Adjust its learned numbers to improve predictions.
The outputs have the same meaning: how likely is each next token? The two models can assign different probabilities.
Either way, generation chooses a next token from those probabilities, appends it, and repeats. We can choose the most likely token or sample, just as in lesson 1.
Why the network can help with new phrases
The important change is what a training example updates. In our count model, it adds evidence to one exact-context row. In the network, it adjusts a shared function used for many different contexts. An update can therefore help with contexts that never appeared exactly in the training text.
That function can learn aspects of meaning and usage: a crash often involves failure; a typo involves text; “fixed” can change the situation. Learning these features lets evidence from one sentence help with another arrangement of familiar words.
The network starts with a learned representation for each token, called an embedding. It encodes useful features as a vector, a list of numbers. The numbers and the network that uses them learn together from next-token prediction; they are not hand-written definitions or labeled features like “severity = 3.”
Let earlier clues change the prediction
We have changed how the model learns from examples. Now we need to give it more context: a network that receives only “build is” still cannot use the earlier “crash.”
An embedding alone cannot tell us what “is” follows here. A transformer is a neural network that updates token representations using the context. Its attention mechanism compares representations across allowed positions and computes weights for mixing information from those positions. Position information lets it account for word order.
In this next-token model, a position can use itself and earlier positions, never the future answer. Repeated attention and further processing of each position build representations that reflect how the words relate in this sentence.
1 · Each token gets a starting representation
Further neural processing refines each representation. These steps repeat through the network.
Which tokens supply information to “is”?
Read the heat map as information flowing into the selected position. Darker cells have larger mixing weights. This shows one attention head: one of the network’s learned ways of combining information across positions.
Illustrative weights. These percentages are chosen to explain the display, not measured from a trained model. Words still stand in for tokens.
Information comes from these positions:
- Position 1the5%
- Position 2crash50%
- Position 3in5%
- Position 4the5%
- Position 5new5%
- Position 6release20%
- Position 7is10%
For “is” at position 7, this illustration gives the largest mixing weight (50%) to information from “crash” at position 2. The allowed weights total 100%. This is one way the earlier issue can reach the position used to predict the next token.
These weights mix information from input positions. They are not probabilities of the next token, or a complete explanation of the model’s answer. Real transformers combine many heads across multiple layers.
Try position 2 to see later tokens blocked. Switch “crash” to “typo” above to change the information being mixed; this illustration keeps the weights fixed so you can separate the information from its mixing weight.
See every position’s attention together
Each row receives information from the columns. Read across a row: its weights total 100%. × marks a later position that is blocked.
Where do these relationships come from?
Our table learned from the tokens that followed each context in the training text. The transformer learns from those same next-token examples, including how to compute useful attention patterns. Training adjusts the network’s numbers so the information it gathers helps its predictions.
It does not store a fixed “is → crash” percentage. After training, the learned numbers stay fixed during ordinary use, but they compute attention weights from the representations in the current sentence. New context can produce a new heat map.
2 · Each position now has a representation informed by its context
An earlier “crash” can affect the representation at the final “is.”
That context-dependent representation can affect the score for “urgent.” The network computes scores for vocabulary tokens, then converts them into probabilities totaling 100%.
Conceptual illustration of a model trained on varied text, not a neural model trained on these four lines. The highlighted route is a possible information path, not measured attention or a guaranteed answer.
The starting embedding of “is” stays the same when you switch inputs. Its processed representation can change. The model keeps a representation at every position; the final one gathers information used for the next prediction. It does not first compress the entire passage into one permanent summary.
In a new phrase, the same learned network combines these features again. It needs no count-table row for that phrase. During ordinary use its learned numbers stay fixed; the representations computed from the current input change.
Quick check
What lets a transformer make an informed guess for a phrase it has never seen exactly?
Further reading and a connection to retrieval
Representation learning also powers query/document embedding models used in retrieval. Those encode whole texts for matching; one token's starting embedding is a different object.
Explore 3Blue1Brown's attention explanation, the original neural language-model paper, and Dive into Deep Learning on attention and context cost. To inspect actual learned numbers in a small Python model, try the optional fixed-window neural experiment. It isolates learned representations using only two context tokens; attention is outside its scope.
Your turn
Explain it in your own words.
Using the four sentences above, the table gives the same 50/50 result for “the crash in the new release is” and “the typo in the new release is.” Why doesn’t matching the whole prefix fix this? How could a trained transformer use the earlier “typo” to make an informed next-token guess for this new phrase?
Answer the question in your own words. A short explanation is enough.
Your feedback
Work through a hint
Hint 1
What observation would the table need for a whole-prefix key? Is that wording in any of the four lines?
Hint 2
What can a shared neural predictor reuse from other sentences, and how can the earlier typo affect the information used to predict after is?
A worked explanation
The whole prefix has no matching observations. The trained network can reuse patterns learned elsewhere and combine them with typo in the current input, changing its next-token scores without needing a row for this exact phrase.