Lesson 6 of 6 · The Shakespeare project

Generate text and inspect your trained model

Your saved checkpoint is a prediction function you can call again. Changing the prompt or the way you choose tokens can change its output without changing its learned numbers. We will separate those changes and inspect real attention from this model.

Reload, predict, append, repeat

model, tokenizer, step = load_checkpoint("my-run/step-4000.pt")
text = generate(model, tokenizer, prompt="\nROMEO:\n", new_tokens=160)

At each step, the program feeds in the last 64 characters, reads the final score row, converts it into probabilities, samples a character, and appends it. The learned numbers stay fixed. A response can grow longer than 64 characters, but each new prediction has access only to the most recent 64.

This lab stops because the application requested 160 new characters. There is no learned EOS token in its continuous-stream vocabulary. A prompt is required; empty input is rejected rather than silently inventing a start token.

Change the sampling, keep the model

Temperature changes how concentrated the sampling probabilities are. At 1, sampling uses the unmodified softmax probabilities from lesson 3. Below 1, high-scoring characters get a larger share; above 1, the distribution spreads out. It changes how we choose from the model’s scores, not what the model has learned.

Recorded results from the downloadable Python project. These controls replay saved measurements; they do not train or run a model in your browser.

ROMEO:
OF QOEXEN EDWARD:
'Tis and a last fast depass yoursess! even his
deple's a seld-tightly, and a such I sweet soon
As it their had belp's drongers; and of it.

Ma

Checkpoint 4000, seed 2026, 160 new characters, temperature 1. Only the selected prompt and temperature vary.

Changing the prompt changes the scores computed from that context. Changing temperature reshapes the sampling probabilities. Choosing a different trained checkpoint would change the learned numbers. These are three different operations.

Read an attention map from the trained model

Lesson 2 used a hand-picked heat map to explain the idea. These maps are measured from our checkpoints. Each position computes comparison scores against allowed positions; softmax turns those scores into mixing weights. The learned comparisons stay fixed during use, but a different input can produce different weights.

A layer here is one transformer block. A head is one of that block’s four parallel attention computations. The grid shows a single head in a single layer. A row receives information from the columns; later columns are blocked.

Measured attention · percentages rounded to one decimal
Receives ↓
Source →
1
⏎ newline
2
R
3
O
4
M
5
E
6
O
7
:
8
⏎ newline
1 · ⏎ newline100.0%×××××××
2 · R90.7%9.3%××××××
3 · O12.4%72.1%15.5%×××××
4 · M1.1%3.8%87.7%7.4%××××
5 · E2.6%0.5%33.2%56.7%6.9%×××
6 · O7.0%1.0%1.2%15.3%54.4%21.2%××
7 · :15.6%22.2%1.7%1.5%15.5%37.9%5.6%×
8 · ⏎ newline0.3%4.4%11.7%0.7%0.9%16.9%57.8%7.3%

Checkpoint 4000 · layer 1 · head 1 · ⏎ROMEO:⏎

Try comparing the untrained and trained checkpoints with the input, layer, and head unchanged. Then inspect another head. Every row’s full-precision weights sum to one, and future positions remain blocked. A darker cell means a larger mixing weight in this head; it is not the probability of the next character.

You may see a pattern you can describe, or a head that is hard to interpret. Either is possible in this small character model. One attention map omits other heads and the rest of the network, so it cannot by itself explain why the model produced a particular word.

Open the attention calculation

The Attention class computes comparison scores, blocks future positions, normalizes the scores, then mixes the value vectors. The optional attention reference explains those vectors and the serving cache; neither is needed to run this project.

Download the complete Python project →

Source, corpus, tests, recorded results, and reference checkpoints. You can follow every example on this page without installing Python.

Run this step on your computer

Unzip the project. Open a terminal in its shakespeare-lab folder. Use Python 3.11–3.13 and create an environment once:

python3 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt

On Windows, see the README for the equivalent commands. For a smaller Linux CPU-only installation, use the command in the README. This lab uses CPU and needs no API key.

python generate.py --checkpoint my-run/step-4000.pt --temperature 0.5
python generate.py --checkpoint my-run/step-4000.pt --inspect

The bundle includes checkpoints from our run so you can inspect the trained model immediately. Your own run can produce different text and slightly different numbers across machines.

Quick check

You reload the same checkpoint and change temperature. What changes?

You have trained a transformer from random initialization, checked it on held-out text, saved it, and used it again. It has learned a small character-level continuation task. Following instructions or answering questions would require training data and objectives that this Shakespeare project does not provide.

Your turn

Explain it in your own words.

Not checked

You use the same saved checkpoint with temperature 0.5 and 1.2. How does that change the probabilities used to choose characters? Does one dark attention cell explain the whole continuation? Explain why.

Answer the question in your own words. A short explanation is enough.

Draft saves on this device0/800 characters

Work through a hint
Code, data, and further reading

This project uses Tiny Shakespeare, public-domain Shakespeare excerpts. Its character-level setup follows the teaching approach in nanoGPT. See PyTorch’s training-loop explanation and its loss API. Our project README documents the exact model, corpus hash, commands and deliberate simplifications.