Lesson 3 of 5 · Fine-tuning for SQL

Fine-tune with a small adapter

We can teach a specific response pattern without updating every number in the original model. A small saved adapter records the change.

SFT describes the learning task; LoRA describes what changes

SFT trains on demonstrated responses. It does not require a particular choice of trainable weights. LoRA, or low-rank adaptation, adds small trainable corrections to selected linear transformations while leaving the pretrained weights frozen. We use SFT with LoRA to make the local experiment manageable.

A linear layer combines input features using a matrix of weights. LoRA adds a second path through two smaller matrices; its output is added to the original layer's output. Those smaller matrices learn the task-specific correction. The optional LoRA paper gives the matrix derivation.

Input featuresFrozen base transformation
+ trainable correction
Output features

Frozen means excluded from optimizer updates. The base model still performs the forward calculation, uses memory, and carries the computation through which the adapter gradients are calculated.

What changed in this run

PartCountDuring training
Original parameters494,032,768Used, but frozen
Adapter parameters2,932,736Updated

The adapters contain about 0.59% as many parameters as the original model. Rank 8 controls the width of their smaller internal path. This run adds adapters to the attention and feed-forward linear projections in the last 16 transformer blocks. It uses the original 16-bit base weights, not a quantized model.

The adapter is about 12 MB on disk. That is not the job's RAM requirement: training also needs the base weights, intermediate computations, gradients, and optimizer state.

Identify which action learns

Recorded experiment · These controls inspect saved outputs or illustrate the procedure. They do not run a model or train in your browser.

Changing a prompt changes the input and may change the answer. It does not update the stored base weights or adapter weights.

Save enough to use the correction again

The adapter weights and configuration must be paired with the same base-model revision, tokenizer, and prompt format. The adapter by itself cannot produce an answer. Our download includes source and recorded outputs; its fetch script downloads the pinned public base model. Your training run creates your own adapter.

Loading an adapter for inference is different from resuming a training job exactly. An exact resume also needs the optimizer's running state and other training state; the inference adapter in this lab does not contain them.

Keep the first experiment small

Before a full run, a short calibration can reveal excessive memory use or broken data. It does not establish that the model learned the task. Run one local model process at a time and keep room for your other applications. If the memory guard refuses to start, use the recorded examples rather than bypassing it.

More rank, more layers, or a larger base model are possible later experiments. Change one factor at a time, compare on development data, and account for their added runtime and memory. You do not need to derive attention or backpropagation to inspect this first pipeline.

Your turn

Explain it in your own words.

Not checked

Only the adapter parameters receive optimizer updates. Why does the frozen base model still need to be loaded, and why can’t the saved 12 MB adapter answer questions by itself?

Answer the question in your own words. A short explanation is enough.

Draft saves on this device0/800 characters

Work through a hint

Sources and further reading

WikiSQL dataset and evaluation rules · MLX-LM training guide · Training, validation, and test sets · Our source and experiment record