Class 2 · Executable examples

Run the calculations in Python

The first three NumPy scripts cover the ML fundamentals. The shared download also retains two language-model examples, now part of the separate Building Language Models from Scratch track. They run on your own computer, not in this browser. Each prints its results and generates a chart.

Download the Python bundle ↓ Read the setup guide →

Download, unzip, and open the folder

Download the bundle above, extract it, and open a terminal in the folder containing requirements.txt and run_all.py. With Python 3 installed, run these commands on macOS or Linux. On Windows, activate the environment with .venv\Scripts\activate instead.

python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
python run_all.py
python -m unittest -v test_examples.py

Run an individual file, such as python 01_single_neuron.py, to focus on one example. Results are also saved under results/ and charts under charts/. These are small CPU examples; no GPU or API key is needed.

Choose a calculation

Sigmoid curve marking a neuron before and after one update
01

One binary neuron

Compute z, sigmoid probability, binary cross-entropy, every gradient, and one SGD update for the example ticket.

Open the Python →
Loss curve with a local tangent, one gradient step, and the plotted optimum
02

Loss and derivative

Compare the local rate of change with the actual optimum, apply the update rule, and recompute the real loss.

Open the Python →
Bar charts converting three logits into softmax probabilities
03

Batch and stable softmax

Execute X @ W + b, inspect each array’s shape, normalize three class scores, and calculate batch gradients.

Open the Python →

Continue with language models

The count-based and neural n-gram examples now belong to Building Language Models from Scratch. Their files remain in the shared bundle for compatibility.

Check your results

Neuronp moves from 0.668 to 0.683 and loss falls from 0.403 to 0.382.
SlopeAt w = −2, dL/dw = −0.917; one step moves to w = −0.166.
SoftmaxLogits [2.0, 0.5, −1.0] become probabilities [0.786, 0.175, 0.039].
N-gramAfter “bug is,” counts 2/1/1 become probabilities 0.50/0.25/0.25.