Lesson 3 of 6 · The Shakespeare project
Make your first transformer prediction
A transformer can run before it has learned anything. This one starts with random numbers, so its first output is mostly gibberish. Follow one input through the working model, then preview the same program after training.
Before training · actual output including the prompt “⏎ROMEO:⏎”
ROMEO: OF3QO rGlWG3WvKwMZ':bALSzlzJdAF-HhH.fl T qaELgYR-3SAXiHMeg'cykNW dUpOTPzkaYOjzStKzgAvwruqy XPaEOncFgISswFYqyRY- -leN3Giu,Y'smkagF&;vKGMLzPr'InUvZJvC&lJX$ i.U Ml
Make the vocabulary small enough to inspect
We will train on Tiny Shakespeare, a collection of Shakespeare excerpts. Here, one character is one token: letters, spaces, punctuation, and newlines each get an ID. That gives us 65 possible next tokens. We preserve capitalization and punctuation.
Lesson 1 used words as tokens. Character tokens let this small model assemble words one character at a time, with a small output table. Modern text LLMs commonly use subword tokens; changing the tokenizer changes the units, not the next-token task.
Our example input is ⏎ROMEO:⏎. The symbol ⏎ marks a newline, so this is eight characters. Each character becomes an integer that selects its embedding. The integer is an address, not a measure of meaning.
Follow the arrays through the model
PyTorch calls its arrays tensors. For now, follow one input. There is one row per character position; columns hold the numbers representing that position. We chose 64 numbers per representation; this width is separate from the 65-character vocabulary.
- 8 token IDsIdentify the eight input characters.
- 8 × 64 numbersLook up a 64-number token embedding and add a learned vector for its position.
- 8 × 64 numbersTwo transformer blocks mix earlier information and process each position’s features.
- 8 × 65 scoresThe output layer rates every possible next character at every position.
The model’s trainable numbers are called parameters, often also called weights. Each block keeps eight positions. A position may use itself and earlier positions, never later ones. Each block runs four attention heads side by side, each with its own learned comparisons, and combines their results. The second block can build on information combined by the first.
Every row predicts the character after its own position. To continue our input, generation uses only the last row: the 65 scores after the final newline. The other rows will become useful when we train on many next-character targets at once.
See the model code and what a block contains
This is the core flow in model.py. The full implementation includes input checks and optional attention recording. token_ids has one extra outer dimension for the number of inputs; here that number is 1.
positions = torch.arange(token_ids.shape[1])
x = self.token_embedding(token_ids) + self.position_embedding(positions)
for block in self.blocks:
x, weights = block(x)
logits = self.output(self.final_norm(x))Inside a block, attention mixes information across positions. A small feed-forward network (MLP) then transforms each position’s features. Adding each result back to its input gives information and training signals a direct path through the block. Layer normalization controls the scale of the numbers. You can follow the prediction without deriving these operations.
The downloadable tests change later characters and verify that earlier output rows do not change. That is a concrete check of the causal rule from lesson 2.
Turn the last row into a next-character guess
Raw scores, also called logits, can be negative and do not have to sum to one. Softmax converts all 65 scores into positive probability shares totaling 100%; larger scores receive larger shares. Compare the score and probability columns below.
Recorded results from the downloadable Python project. These controls replay saved measurements; they do not train or run a model in your browser.
8 input positions → 8 rows of 65 scores → use row 8 to continue
| Character | Raw score | Probability |
|---|---|---|
| . | 0.417 | 2.30% |
| z | 0.272 | 1.99% |
| A | 0.240 | 1.93% |
| 3 | 0.215 | 1.88% |
| ⏎ newline | 0.202 | 1.85% |
| q | 0.188 | 1.83% |
The other 59 characters share 88.22%. Displayed numbers are rounded.
Sampling chooses a character using these shares, just as lesson 1 sampled a word. Append that character and run the predictor again. Softmax and sampling use the current learned numbers; neither operation trains the model.
ROMEO: OF3QO rGlWG3WvKwMZ':bALSzlzJdAF-HhH.fl T qaELgYR-3SAXiHMeg'cykNW dUpOTPzkaYOjzStKzgAvwruqy XPaEOncFgISswFYqyRY- -leN3Giu,Y'smkagF&;vKGMLzPr'InUvZJvC&lJX$ i.U Ml
Random weights · prompt ⏎ROMEO:⏎ · seed 2026 · 160 new characters. Samples use the unmodified softmax shares (temperature 1, explained in lesson 6). The trained preview keeps these settings.
Source, corpus, tests, recorded results, and reference checkpoints. You can follow every example on this page without installing Python.
Run this step on your computer
Unzip the project. Open a terminal in its shakespeare-lab folder. Use Python 3.11–3.13 and create an environment once:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txtOn Windows, see the README for the equivalent commands. For a smaller Linux CPU-only installation, use the command in the README. This lab uses CPU and needs no API key.
python generate.py
python generate.py --checkpoint reference/step-4000.ptThe bundle includes checkpoints from our run so you can inspect the trained model immediately. Your own run can produce different text and slightly different numbers across machines.
Quick check
Why can the untrained model return valid probabilities and still produce nonsense?
The trained preview uses the same program with different learned numbers. Next, we will use the character that actually followed a context to show how one update changes those numbers.
Your turn
Explain it in your own words.
The model receives eight characters and returns eight rows of 65 scores. Which row do we use to continue the input, how do its scores become next-character probabilities, and does choosing a character update the learned model numbers?
Answer the question in your own words. A short explanation is enough.
Your feedback
Work through a hint
Hint 1
Which position comes immediately before the character we want to add?
Hint 2
Are raw scores already percentages? Is this program applying a training update?
A worked explanation
Use the last row, because it predicts what comes after the whole supplied input. Turn its scores into probability shares and choose from those. The model’s learned numbers do not change.
Code, data, and further reading
This project uses Tiny Shakespeare, public-domain Shakespeare excerpts. Its character-level setup follows the teaching approach in nanoGPT. See PyTorch’s training-loop explanation and its loss API. Our project README documents the exact model, corpus hash, commands and deliberate simplifications.