# Train a tiny transformer on Shakespeare

This is the runnable project for lessons 3–6 of Building Language Models from Scratch:
https://dougdoes.ai/courses/llms-from-first-principles/tracks/language-models-in-python/inside.html

Start with a random model, observe one update, train it, then reload a checkpoint
and inspect its predictions and attention. Everything runs on CPU. No API keys,
GPU, pretrained download, or paid service is needed.

## Setup

Use Python 3.11–3.13. Unzip the download and open a terminal in `shakespeare-lab`.
On macOS or Linux:

```sh
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt
```

On Windows PowerShell:

```powershell
py -m venv .venv
.venv\Scripts\python -m pip install -r requirements.txt
.venv\Scripts\python generate.py
```

The explicit Python path avoids needing to change PowerShell's script policy.
Use it in place of `python` in the commands below. On Linux you can install the
smaller CPU-only PyTorch distribution:

```sh
python -m pip install torch==2.8.0 --index-url https://download.pytorch.org/whl/cpu
python -m pip install 'numpy>=1.26,<3'
```

The published run used Python 3.12.14, PyTorch 2.8.0, NumPy 2.5.3, macOS arm64,
and two CPU threads. It took 31 seconds, with approximately 303 MiB peak process
memory. Installation uses more disk space than the model; allow several GB.
Runtime and numerical results can vary across machines. Seeds make comparisons
repeatable within a compatible environment, not bit-for-bit portable everywhere.

## Follow the lessons

```sh
# Lesson 3: generate with random weights, then a supplied trained checkpoint.
python generate.py
python generate.py --checkpoint reference/step-4000.pt

# Lesson 4: one update on the first 64 targets. Prints before/after measurements.
python one_update.py

# Lesson 5: a fresh training run. Choose an output folder that does not exist.
python train.py --steps 4000 --out my-run

# Lesson 6: reload your model. This does not train again.
python generate.py --checkpoint my-run/step-4000.pt --temperature 0.5
python generate.py --checkpoint my-run/step-4000.pt --prompt 'JULIET:'
python generate.py --checkpoint my-run/step-4000.pt --inspect

# Check causality, shifted targets, update behavior, and checkpoint reloading.
python -m unittest -v
```

For a quick installation check use `--steps 50 --out smoke-run`; 50 steps will not
reproduce the fully trained sample. To regenerate all lesson evidence after a
full run, including the separate one-window overfitting demonstration:

```sh
python export_examples.py my-run --output my-results.json
```

`results.json` contains the published run. The download also includes reference
checkpoints at 0, 500, 2,000 and 4,000 updates, so you can explore without first
training. `--inspect` prints full-precision attention weights for every layer and
head, plus final-position scores and probabilities. If the prompt is longer than
the model's window, only its last 64 characters are inspected.

## What this model does

- 65 character tokens: preserve spaces, newlines, punctuation, and capitalization.
- 64-character context window, 64 features per position, 2 blocks, 4 heads per block.
- 112,577 trainable numbers. Token and position embeddings are learned together
  with attention, feed-forward networks, normalization, and the output layer.
- Causal attention prevents each position from reading later input positions.
- Each 65-character slice supplies 64 input/next-character pairs.
- Batches contain 16 windows. AdamW uses learning rate 0.001 and weight decay 0.01.
- Model initialization and training-batch seeds are 1337. Evaluation seed is 9001.
  Published samples use seed 2026, 160 new characters, and temperature 1.0 unless
  a comparison explicitly specifies another temperature.

We split the raw file first: the first 1,003,854 characters are training text and
the final 111,540 validation text. Windows never cross this split. The curve
scores the same 128 sampled windows per split at each measurement; validation
never contributes parameter updates. This is not a separate final test set for
selecting among many experiments.

## Deliberate simplifications

This is a continuous stream, **without BOS or EOS tokens**. Every generation
requires a nonempty prompt, and the application stops at its requested character
limit, possibly mid-word. Newlines are ordinary characters, not hidden sequence
markers. Unlike lesson 1's separate sequences, a training window is not a new
document, and no EOS target is added to it. A production dataset can choose
explicit document boundaries and train their tokens as targets.

We recompute the most recent context on every generated character; this small
teaching implementation has no KV cache. Context is limited to 64 characters.
There is no dropout, instruction tuning, retrieval system, or chat interface.
It learns continuation patterns, not reliable answers to questions.

Checkpoints contain weights, model configuration, token mapping, and corpus
hash. They support generation and inspection, **not exact interrupted-training
resume**, because optimizer and training-generator state are not saved.

## Data and source

`input.txt` is Tiny Shakespeare, a roughly 1 MB selection of public-domain
Shakespeare text, from Andrej Karpathy's char-rnn repository:
https://github.com/karpathy/char-rnn/tree/master/data/tinyshakespeare

Corpus SHA256:
`86c4e6aa9db7c042ec79f339dcb96d42b0075e16b8fc2e86bf0ca57e2dc565ed`

`load_data()` checks that hash. This implementation follows the small
character-transformer teaching approach of https://github.com/karpathy/nanoGPT.
The model is implemented here for this course; each Python file is included.
