By Doug Schonholtz

A practical study companion to Karpathy’s Zero to Hero

Pair the full code-along course with small experiments, explanations, and checkpoints between sessions.

Last checked: October 3, 2026.

Keep the implementation work

If your goal is to understand neural networks by building them, make Neural Networks: Zero to Hero your main course. A small interactive lesson can help you investigate a confusing step or explain it back. Keep the full implementation sequence as the main work.

Bring solid Python and introductory mathematics. Expect to pause, write code, inspect mistakes, and repeat sections. The official course repository supplies notebooks and points to exercises in the video descriptions. Watching a solution and independently rebuilding it are different tasks.

Know what you are budgeting for

These are the video lengths listed in the official syllabus, rounded to minutes. They exclude setup, exercises, debugging, and time spent trying your own variations. The syllabus is marked ongoing; this table is the eight entries it lists at the October 3 check, not every video Karpathy has published.

Syllabus topic Listed viewing time
micrograd and backpropagation 2h 25m
makemore: character bigrams 1h 57m
makemore: multilayer perceptrons 1h 15m
Activations, gradients, and BatchNorm 1h 55m
Manual backpropagation through the MLP 1h 55m
WaveNet-style architecture 56m
Building GPT 1h 56m
Building a GPT tokenizer 2h 13m

Source: Karpathy’s syllabus and video links. The displayed lengths total about 14 hours 32 minutes of video. Your study time will be longer if you do the practical work.

Four checkpoints to use alongside the course

The checkpoints below are suggested exercises for this study plan. They are not graded prerequisites for Karpathy’s course.

1. Can you explain one gradient before building an engine?

Use Doug Does AI’s gradient lesson for a deliberately small calculation. Predict what happens to the loss when one weight changes slightly, then check it. Explain the sign of the gradient and the direction of an update.

Then keep going with Karpathy’s micrograd implementation. The micrograd repository includes a tiny scalar autodiff engine, a neural-network example, and tests that compare gradients with PyTorch. Building and testing that engine goes beyond the small browser calculation.

A useful checkpoint: choose a small expression, calculate a derivative, estimate it with a tiny numerical change, and compare both with the implementation. If the values disagree, investigate before adding a larger network.

2. Can you separate counts, scores, probabilities, and samples?

Pair the makemore work with Doug Does AI’s count-based language-model lesson and softmax module. First inspect observed next-token counts. Then distinguish a model’s output scores from the probabilities used to choose a continuation.

The makemore project works with character-level sequences and offers several model architectures. Doug Does AI’s count lesson begins with word-level examples. The token units differ, so do not compare their losses as if they were the same experiment.

A useful checkpoint: explain what happens when a context has no observations, why a shared learned model can behave differently from a lookup table, and why a plausible generated sample is not enough to establish generalization.

3. Can you trace a small transformer without treating it as a chatbot?

Use Doug Does AI’s Shakespeare track as a second compact example: follow a prediction, make an update, train, reload a checkpoint, and inspect generation. It includes a CPU Python project. Prepared browser examples let you examine recorded results without installing the project.

In the inspection lesson, changing temperature changes the sampling distribution while the saved model stays fixed. The measured attention maps show individual heads, not a complete explanation of an answer.

A useful checkpoint: explain the difference between changing the prompt, changing temperature, and training another checkpoint. Identify which operation changes learned parameters. Then explain why a character-continuation model trained on Shakespeare has not thereby learned to follow instructions.

Return to Karpathy’s GPT implementation to connect those observations to the code you are building.

4. Can you identify the parts this companion does not replace?

Keep the deeper makemore sections when you want practice with network diagnostics, BatchNorm, manual tensor backpropagation, and the WaveNet-style hierarchy. Those topics do not have matching full lessons in Doug Does AI’s current core tracks. A six-lesson Shakespeare route is not evidence of equivalent coverage.

Doug Does AI’s Shakespeare project also uses a character vocabulary. For tokenizer implementation, use Karpathy’s tokenizer lecture and minbpe, which includes a progression exercise and tests. A character lookup does not give you the same implementation experience as training and using a byte-pair tokenizer.

A useful checkpoint: explain why changing a tokenizer changes the token sequence the model sees. Encode and decode a small example, check that the text round-trips, and distinguish tokenizer training from updating the language model’s parameters.

A 25-minute session you can repeat

This is a scheduling suggestion, not a completion-time claim.

  1. Five minutes: without replaying anything, write what you remember and name one uncertainty.
  2. Fifteen minutes: work on one bounded piece. That might be a video section, a derivative check, a failing test, or one interactive example. Stop at a meaningful boundary rather than forcing a full lesson into the timer.
  3. Five minutes: explain what changed, save the code or result, and write the next question to investigate.

Reserve longer blocks for environment setup and debugging. A long video can be split across days; a shorter page can still require substantial practice.

How to use feedback

Try to predict an outcome before checking a solution. Runnable code, numerical checks, and the original reference implementation give you concrete evidence to compare.

Doug Does AI’s public lessons include written hints and worked explanations. AI feedback on your own explanation and page chat require sign-in, eligible access, and available usage. Treat that feedback as assistance: it can make mistakes, and it does not replace tests or establish mastery. Recorded browser model examples do not run training in the browser.

If the code disagrees with an explanation, reduce the example until you can inspect each step. If you ask a community for help, include the smallest reproducible case and what you expected to happen.

Where to go after this

If you want to adapt a pretrained model rather than train a small one from scratch, Hugging Face’s fine-tuning chapter introduces practical training workflows. Doug Does AI’s SQL fine-tuning track provides a narrower example of SFT, LoRA, and evaluation. That is a different stage of work, not a missing chapter silently attributed to Zero to Hero.

For broader choices, see the AI learning guide. Pick the resource that matches the next thing you want to be able to explain or build.