The technical track · Basic Python; no ML background required

Train a small transformer on Shakespeare

Start with a word-count predictor you can inspect. See why it struggles with new phrases and distant clues, then train a neural network that learns patterns from text.

Six lessons take you from counts to a working character-level transformer. Each page includes examples you can explore without installing Python. The downloadable project runs on CPU and includes source, data, tests, and trained checkpoints. The optional ML fundamentals modules offer more mathematics.

Start with counts →Download the transformer project

From observed counts to learned predictions

Lesson 1

A language model made of counts

Build next-word probabilities from observations, then find where unseen contexts defeat the table.

Open lesson →
Lesson 2

How transformers use context to predict new text

See how learned representations and attention use earlier clues in phrases the model has never seen.

Open lesson →
Lesson 3

Make your first transformer prediction

Follow Shakespeare characters through a working model and turn its output scores into probabilities.

Open lesson →
Lesson 4

Teach the model with one update

Use the observed next characters to score a prediction and adjust the learned numbers.

Open lesson →
Lesson 5

Train your transformer on Shakespeare

Repeat updates across many windows, check held-out text, and save the resulting model.

Open lesson →
Lesson 6

Generate text and inspect your trained model

Reload your checkpoint, change sampling settings, and inspect its actual attention maps.

Open lesson →

Next: adapt a pretrained model

Continue with Fine-tuning a Model for SQL: use demonstrated responses to train a small adapter, then check whether the answers improve.

Learning guides

Studying alongside a full implementation course? Use the Karpathy Zero to Hero study companion for four checkpoints and a plan for the video sequence.

For other routes, compare six AI learning resources by goal, prerequisites and practice.

Optional practice and references

Shakespeare with unigrams, bigrams, and trigrams extends lesson 1 using only Python’s standard library. The attention reference explains the vector calculation and generation cache after lesson 6.

For a smaller neural model without attention, the NumPy practice bundle includes count-based and neural n-grams. Follow its setup instructions.