Optional · Executable neural n-gram baseline

A small fixed-window neural baseline

This optional Python experiment isolates shared embeddings and a next-token predictor. It uses only two context tokens and has no attention: it is a neural n-gram model, not a transformer. Return to the main lesson.

1 · Pick up where counting stopped

What could follow “issue is”?

Our new training text contains “bug is,” “bug was,” and “issue was,” but never “issue is.” Both words are known. The exact-count model has no matching row, so the version we built in lesson 1 returns no estimate.

An embedding is a learned representation of a token's meaning and how it is used. Here, “bug” and “issue” both describe a problem that can be urgent or reproducible. We want the model to capture that shared meaning so that what it learns about one word can help it make predictions with the other.

The model learns that relationship from examples: both words appear before “was urgent” and “was reproducible.” It stores its representation as a vector—a list of numbers. Training adjusts those numbers and the network that uses them, turning patterns in the text into knowledge it can reuse when it encounters a new combination such as “issue is.”

Follow the prediction

From a token's address to a learned guess

These are recorded snapshots from the downloadable Python program. Switch between its random starting values and its trained values, or try a context that did appear.

All the training text
the bug is urgent
the bug is reproducible
the bug was urgent
the bug was reproducible
the issue was urgent
the issue was reproducible
the build is broken
the build was broken

Each line is a separate sequence. We keep lesson 1's two start markers and one end target.

The exact-count lookup

  1. 1 · Token IDs

    Find the addresses

    IDs select rows. Their numeric distance does not tell us which words are related.

  2. 2 · Embedding lookup

    Look up the learned representation

    This row encodes patterns learned about the token's meaning and use. The same token reuses this representation in different contexts.

  3. 3 · Shared neural layer

    Calculate a score per output

    8 embedding values feed 10 output scoresTwo token rows supply four numbers each. Every input connects to every output score through a trainable weight.8 input values10 scores

    Each line is a learned connection. The same connections combine the input numbers for every context.

  4. 4 · Softmax

    Turn scores into probabilities

    Softmax gives larger scores larger shares, with all output probabilities adding to 100%.

This drawing shows our tiny lookup-plus-linear-layer network. Bias values are included in its calculations but omitted from the drawing. Modern language models use much richer context-processing networks.

Every possible next token
Next tokenRaw scoreProbability after softmax

2 · Where the relationships come from

Train shared parts instead of a separate row for each phrase

We do not tell the model that “bug” and “issue” are related. In this corpus, both appear before “was urgent” and “was reproducible.” Training on these similar roles can make their representations produce similar next-token predictions.

To predict after “issue is,” the network retrieves the issue row learned in other contexts and the is row learned with other words. It combines them using the same learned connections. No exact “issue is” row is required.

The computed result

After 800 training steps on 40 context/target rows, the unseen issue is receives 49.0% for urgent and 49.0% for reproducible. Other outputs share the remaining probability. The exact-count table still has no row for this context.

This deliberately small example illustrates transfer to one missing combination. It does not establish broad language understanding or guarantee that every new combination gets a good prediction.

One training example, step by step

  1. Read the context from the training text, such as bug is.
  2. Predict the next token and compare the probabilities with what actually followed, such as urgent.
  3. Calculate adjustments that favor the observed target. Training applies small updates to the relevant embedding rows and the shared connections.
  4. Repeat across examples. Later predictions reuse those updated numbers, including in contexts not seen during training.

The Python program calculates these adjustments for every training row, averages them, then updates the model once per step.

The starting rows are random; their useful structure comes from training. Each coordinate is a learned feature, not a named field such as “is a bug.” Relationships can be spread across several coordinates.

Quick check 1

If bug has ID 1 and issue has ID 4, where can their useful relationship be learned?

3 · Read the output

A score is not yet a probability

The network produces a raw score for each possible output. Scores can be negative and need not add to anything in particular. Softmax converts them into positive shares that add to one. Higher-scoring tokens get larger shares.

Just the conversion · three illustrative scores

For raw scores [2, 1, 0], softmax produces about 66.5%, 24.5%, 9.0%. Scores differ by one, but these probabilities do not.

Softmax uses a positive weight derived from each score, then divides by the total weight. It does not count training occurrences or look up word meanings.

The random model in the example also produces probabilities. Training the embeddings and connections is what makes the scores useful. Softmax converts the scores; it does not establish whether the prediction is sensible.

Optional: see the arithmetic and training terminology

Softmax takes exp(score) for each output and divides it by the sum of those values. In code, subtract the largest score first to keep the numbers manageable; this leaves the probabilities unchanged.

weights = exp(scores - max(scores))
probabilities = weights / sum(weights)

The program scores a training prediction with cross-entropy: for one observed target, it is −log(the probability assigned to that target). Giving the actual target more probability lowers this error score. Backpropagation calculates how each trainable number contributes to the error; the optimizer uses those calculations to make updates. You do not need those formulas to explain the sharing mechanism.

Quick check 2

What does softmax contribute to this prediction?

4 · Connect this to the missing-data problem

Make a guess for a combination you have not counted

Lesson 1's smoothing methods adjusted count estimates to leave room for unobserved outcomes. This neural model addresses the same shortage of exact observations through shared learned representations and a shared prediction function. It does not add pretend observations to the count table.

“Unseen” here means an unseen combination of known tokens. An entirely new token ID has no learned row to retrieve. A modern tokenizer can often split an unfamiliar word into familiar pieces, allowing a model to work with those known tokens. This toy word tokenizer cannot do that. A randomly added row would still need training.

The neural language-model paper behind this idea explains how learned word representations support predictions for new sequences.

Python bridge · run the worked example
token_vectors = embeddings[context_token_ids]
scores = token_vectors.reshape(1, -1) @ weights + bias
probabilities = softmax(scores)

Download the Python bundle and run python3 06_embedding_generalization.py. It prints the before/after predictions, including the held-out context. This experiment needs NumPy; it does not download a pretrained model. Setup instructions.

The previous corpus's neural example uses the same training function. The corpus changed here to make the missing-combination question visible.

Optional: array sizes and memory

This experiment uses 11 input IDs: nine ordinary words, BOS, and EOS. Its 10 possible outputs exclude BOS. Each input row holds four numbers; two rows supply eight inputs to the scoring layer. The original count-lesson corpus uses nine input IDs and eight outputs.

Embedding table
Lookup result
Combined context
Output scores

embedding parameters ≈ using two bytes per value. This estimate excludes the scoring layer and the rest of a larger model. The larger vocabulary settings assume equally many input and output token types.