Building Language Models from Scratch · Optional Python practice
Predict Shakespeare’s next word by counting
Can a dictionary of word counts predict what comes next? Build three small models that use different amounts of context. Then see why remembering more words helps—and why counts alone eventually run out of examples.
All lessons in this track · The n-gram lesson
This optional exercise extends lesson 1’s count model to Shakespeare. It needs Python 3.10 or newer and no packages. To train the transformer, continue with lesson 3’s cumulative project.
Step 1
Use a short passage as the training data
The script includes the opening of Hamlet’s “To be, or not to be” speech, ending at “Must give us pause.” It lowercases the text and discards punctuation. Here a token is a whole word; this is not an LLM tokenizer. We treat the passage as one sequence, ignoring line breaks, with start and end markers.
Input text: To be, or not to be: that is the question Words: to | be | or | not | to | be | that | is | the | question Question for all three models: What word comes next after “to be”?
The models learn only from this tiny excerpt, not all of Shakespeare or all of English. A high probability here means a frequent continuation in this data—not a universal rule of language.
Step 2
Change how much history the model uses
| Model | History used | How it estimates the next word |
|---|---|---|
| Unigram | None | P(word) = its count / all predicted tokens |
| Bigram | Last word: be | P(word | be) = count(be, word) / all continuations after be |
| Trigram | Last two words: to be | P(word | to be) = count(to, be, word) / all continuations after to be |
The name counts the predicted word too: a trigram uses two previous words to predict the third. The input can be a long sentence, but a trigram ignores everything before those last two words.
In this excerpt, “be” appears three times, always after “to.” Its next words are “or,” “that,” and “wish’d,” once each. Both the bigram and trigram therefore give those three continuations one-third probability. More context does not automatically change the answer.
Now try “to sleep”: Bigram: P(next word | sleep) to: 40% no: 20% perchance: 20% of: 20% Trigram: P(next word | to sleep) no: 33⅓% to: 33⅓% perchance: 33⅓%
The bigram also counts “a sleep to” and “that sleep of.” The trigram excludes those because their two-word context is different. The unigram ignores both prompts and always returns the same overall word frequencies.
These are distributions over the vocabulary observed in our passage, including an end marker. Unlisted words have zero probability in this deliberately unsmoothed model. Real count-based models often use smoothing or shorter-context fallback to avoid treating unseen continuations as impossible.
Step 3
Run it and change the prompt
Download ngrams.py, then run these commands in its folder:
python3 ngrams.py python3 ngrams.py --prompt "to sleep" python3 ngrams.py --prompt "purple sleep" python3 ngrams.py --prompt "to sleep" --seed 12 # Optional: supply your own UTF-8 plain-text training data python3 ngrams.py --corpus my-text.txt --prompt "to sleep"
The output shows all nonzero next-word probabilities, then a generated continuation for each model. Each sampled word is appended before the next prediction. A fixed random seed makes comparisons repeatable. Generation stops at the end marker, the step limit, or an unseen context; it never silently switches to another model.
With “purple sleep,” the bigram still uses “sleep.” The trigram finds no “purple sleep” context and cannot estimate a distribution. That is missing evidence, not proof that no word can follow it.
Step 4
What actually becomes too large?
A unigram needs one count per vocabulary word. A sparse bigram or trigram stores only combinations seen in the training text; both are practical, including in this script. A table reserving space for every possible combination is a different proposition.
| 50,000-word vocabulary | Possible entries | Four bytes per entry |
|---|---|---|
| Unigram: V | 50,000 | 200 KB |
| Bigram: V² | 2.5 billion | 10 GB |
| Trigram: V³ | 125 trillion | 500 TB |
These are decimal storage units for the probability values alone, not Python dictionary memory estimates. The script prints these totals but does not allocate those tables. Adding another word of context multiplies the possible entries by 50,000 again.
Storage is only half the problem. Even if the table fit, many longer phrases would never appear in the training data. Our “purple sleep” example already runs into that problem with two words of context.
Step 5
Why use a neural network instead?
A neural language model learns shared parameters that estimate next-token probabilities from a context. It does not need a separate stored count for every exact phrase. Training can teach patterns that transfer to combinations it has never seen, while allowing much longer contexts than this toy model uses. That does not guarantee its continuation is sensible or true.
The job stays the same: estimate a distribution, select a token, append it, and repeat. The change is how we estimate that distribution.
Explain the difference
Why does changing “to sleep” to “purple sleep” leave the bigram’s prediction unchanged, but stop the trigram? Before opening the explanation, name the exact words each model looks up.
Compare your reasoning
The bigram uses only “sleep,” so both inputs look up the same counts. The trigram uses two words. It has observations for “to sleep,” but none for “purple sleep.” This implementation has no fallback, so it stops instead of inventing probabilities. It has not discovered that the phrase is impossible.
Further reading: Jurafsky and Martin: N-gram Language Models. The bundled training passage is public-domain Shakespeare, Hamlet, Act III, scene i.