A language model made of counts
Build next-word probabilities from observations, then find where unseen contexts defeat the table.
Open lesson →The technical track · Basic Python; no ML background required
Start with a word-count predictor you can inspect. See why it struggles with new phrases and distant clues, then train a neural network that learns patterns from text.
Six lessons take you from counts to a working character-level transformer. Each page includes examples you can explore without installing Python. The downloadable project runs on CPU and includes source, data, tests, and trained checkpoints. The optional ML fundamentals modules offer more mathematics.
Build next-word probabilities from observations, then find where unseen contexts defeat the table.
Open lesson →See how learned representations and attention use earlier clues in phrases the model has never seen.
Open lesson →Follow Shakespeare characters through a working model and turn its output scores into probabilities.
Open lesson →Use the observed next characters to score a prediction and adjust the learned numbers.
Open lesson →Repeat updates across many windows, check held-out text, and save the resulting model.
Open lesson →Reload your checkpoint, change sampling settings, and inspect its actual attention maps.
Open lesson →Continue with Fine-tuning a Model for SQL: use demonstrated responses to train a small adapter, then check whether the answers improve.
Studying alongside a full implementation course? Use the Karpathy Zero to Hero study companion for four checkpoints and a plan for the video sequence.
For other routes, compare six AI learning resources by goal, prerequisites and practice.
Shakespeare with unigrams, bigrams, and trigrams extends lesson 1 using only Python’s standard library. The attention reference explains the vector calculation and generation cache after lesson 6.
For a smaller neural model without attention, the NumPy practice bundle includes count-based and neural n-grams. Follow its setup instructions.