Optional · Executable neural n-gram baseline
A small fixed-window neural baseline
This optional Python experiment isolates shared embeddings and a next-token predictor. It uses only two context tokens and has no attention: it is a neural n-gram model, not a transformer. Return to the main lesson.
1 · Pick up where counting stopped
What could follow “issue is”?
Our new training text contains “bug is,” “bug was,” and “issue was,” but never “issue is.” Both words are known. The exact-count model has no matching row, so the version we built in lesson 1 returns no estimate.
An embedding is a learned representation of a token's meaning and how it is used. Here, “bug” and “issue” both describe a problem that can be urgent or reproducible. We want the model to capture that shared meaning so that what it learns about one word can help it make predictions with the other.
The model learns that relationship from examples: both words appear before “was urgent” and “was reproducible.” It stores its representation as a vector—a list of numbers. Training adjusts those numbers and the network that uses them, turning patterns in the text into knowledge it can reuse when it encounters a new combination such as “issue is.”
Follow the prediction
From a token's address to a learned guess
These are recorded snapshots from the downloadable Python program. Switch between its random starting values and its trained values, or try a context that did appear.
This drawing shows our tiny lookup-plus-linear-layer network. Bias values are included in its calculations but omitted from the drawing. Modern language models use much richer context-processing networks.
2 · Where the relationships come from
Train shared parts instead of a separate row for each phrase
We do not tell the model that “bug” and “issue” are related. In this corpus, both appear before “was urgent” and “was reproducible.” Training on these similar roles can make their representations produce similar next-token predictions.
To predict after “issue is,” the network retrieves the issue row learned in other contexts and the is row learned with other words. It combines them using the same learned connections. No exact “issue is” row is required.
The computed result
After 800 training steps on 40 context/target rows, the unseen issue is receives 49.0% for urgent and 49.0% for reproducible. Other outputs share the remaining probability. The exact-count table still has no row for this context.
This deliberately small example illustrates transfer to one missing combination. It does not establish broad language understanding or guarantee that every new combination gets a good prediction.
One training example, step by step
The Python program calculates these adjustments for every training row, averages them, then updates the model once per step.
The starting rows are random; their useful structure comes from training. Each coordinate is a learned feature, not a named field such as “is a bug.” Relationships can be spread across several coordinates.
Quick check 1
If bug has ID 1 and issue has ID 4, where can their useful relationship be learned?
3 · Read the output
A score is not yet a probability
The network produces a raw score for each possible output. Scores can be negative and need not add to anything in particular. Softmax converts them into positive shares that add to one. Higher-scoring tokens get larger shares.
Just the conversion · three illustrative scores
For raw scores [2, 1, 0], softmax produces about 66.5%, 24.5%, 9.0%. Scores differ by one, but these probabilities do not.
Softmax uses a positive weight derived from each score, then divides by the total weight. It does not count training occurrences or look up word meanings.
The random model in the example also produces probabilities. Training the embeddings and connections is what makes the scores useful. Softmax converts the scores; it does not establish whether the prediction is sensible.
Quick check 2
What does softmax contribute to this prediction?
4 · Connect this to the missing-data problem
Make a guess for a combination you have not counted
Lesson 1's smoothing methods adjusted count estimates to leave room for unobserved outcomes. This neural model addresses the same shortage of exact observations through shared learned representations and a shared prediction function. It does not add pretend observations to the count table.
“Unseen” here means an unseen combination of known tokens. An entirely new token ID has no learned row to retrieve. A modern tokenizer can often split an unfamiliar word into familiar pieces, allowing a model to work with those known tokens. This toy word tokenizer cannot do that. A randomly added row would still need training.
The neural language-model paper behind this idea explains how learned word representations support predictions for new sequences.
token_vectors = embeddings[context_token_ids]
scores = token_vectors.reshape(1, -1) @ weights + bias
probabilities = softmax(scores)Download the Python bundle and run python3 06_embedding_generalization.py. It prints the before/after predictions, including the held-out context. This experiment needs NumPy; it does not download a pretrained model. Setup instructions.
The previous corpus's neural example uses the same training function. The corpus changed here to make the missing-combination question visible.