# Class 2 Python track

These six examples make the lessons executable. They print numerical results and write machine-readable results under `results/`. The first five can also regenerate the lecture charts.

To run just the count model, download `04_ngram_language_model.py` and run
`python3 04_ngram_language_model.py`. This path needs only Python's standard
library (Python 3.10+). It prints every training row, saves its JSON, and traces
generation from an empty history to the end marker. Add `--chart` after installing
the dependencies below and downloading `course_utils.py` to produce its chart.

The count and neural n-grams now use two `<BOS>` context markers and one `<EOS>`
target per nonempty sequence: five four-word sequences produce 25 rows. Their
ordinary tokens use the same lowercasing and punctuation rules as the lesson.
The neural example has nine input token types (seven words and two markers),
and eight possible outputs (seven words and EOS). BOS is never a target.

The second lesson's experiment is `06_embedding_generalization.py`. It uses a
different, explicitly listed eight-line corpus that never contains `issue is`.
Both words occur elsewhere. Run it with NumPy to compare the random and trained
model's guesses for that unseen combination. Its 40 training rows use 11 input
IDs and 10 possible outputs. `embedding_model.py` supplies the shared training
and prediction functions used by examples 05 and 06; keep this helper beside
those scripts. Example 06 does not need Matplotlib.

```bash
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
python run_all.py
python -m unittest -v test_examples.py
```

The sequence is cumulative:

1. `01_single_neuron.py` — forward pass, binary cross-entropy, manual gradient, and one SGD update.
2. `02_loss_and_gradient.py` — the loss function, its local derivative, the plotted optimum, and a recomputed post-update loss.
3. `03_batch_and_softmax.py` — `X @ W + b`, tensor shapes, stable softmax, cross-entropy, and batch gradients.
4. `04_ngram_language_model.py` — sliding-window training rows, continuation counts, and a conditional next-token distribution.
5. `05_neural_ngram.py` — learned embeddings plus a linear language-model head trained with manual NumPy gradients.
6. `06_embedding_generalization.py` — learned predictions for a held-out combination of known tokens, compared with exact counts.

Everything is deterministic. If a displayed value in the slides changes, rerun `python run_all.py` and rebuild the deck so the code, charts, and lecture stay aligned.

The test suite checks the forward pass, analytical-versus-finite-difference gradients, an improving update, softmax normalization and shift invariance, and the expected trigram distribution.
