Lesson 3 of 5 · Fine-tuning for SQL
Fine-tune with a small adapter
We can teach a specific response pattern without updating every number in the original model. A small saved adapter records the change.
SFT describes the learning task; LoRA describes what changes
SFT trains on demonstrated responses. It does not require a particular choice of trainable weights. LoRA, or low-rank adaptation, adds small trainable corrections to selected linear transformations while leaving the pretrained weights frozen. We use SFT with LoRA to make the local experiment manageable.
A linear layer combines input features using a matrix of weights. LoRA adds a second path through two smaller matrices; its output is added to the original layer's output. Those smaller matrices learn the task-specific correction. The optional LoRA paper gives the matrix derivation.
+ trainable correctionOutput features
Frozen means excluded from optimizer updates. The base model still performs the forward calculation, uses memory, and carries the computation through which the adapter gradients are calculated.
What changed in this run
| Part | Count | During training |
|---|---|---|
| Original parameters | 494,032,768 | Used, but frozen |
| Adapter parameters | 2,932,736 | Updated |
The adapters contain about 0.59% as many parameters as the original model. Rank 8 controls the width of their smaller internal path. This run adds adapters to the attention and feed-forward linear projections in the last 16 transformer blocks. It uses the original 16-bit base weights, not a quantized model.
The adapter is about 12 MB on disk. That is not the job's RAM requirement: training also needs the base weights, intermediate computations, gradients, and optimizer state.
Identify which action learns
Recorded experiment · These controls inspect saved outputs or illustrate the procedure. They do not run a model or train in your browser.
Changing a prompt changes the input and may change the answer. It does not update the stored base weights or adapter weights.
Save enough to use the correction again
The adapter weights and configuration must be paired with the same base-model revision, tokenizer, and prompt format. The adapter by itself cannot produce an answer. Our download includes source and recorded outputs; its fetch script downloads the pinned public base model. Your training run creates your own adapter.
Loading an adapter for inference is different from resuming a training job exactly. An exact resume also needs the optimizer's running state and other training state; the inference adapter in this lab does not contain them.
Keep the first experiment small
Before a full run, a short calibration can reveal excessive memory use or broken data. It does not establish that the model learned the task. Run one local model process at a time and keep room for your other applications. If the memory guard refuses to start, use the recorded examples rather than bypassing it.
More rank, more layers, or a larger base model are possible later experiments. Change one factor at a time, compare on development data, and account for their added runtime and memory. You do not need to derive attention or backpropagation to inspect this first pipeline.
Your turn
Explain it in your own words.
Only the adapter parameters receive optimizer updates. Why does the frozen base model still need to be loaded, and why can’t the saved 12 MB adapter answer questions by itself?
Answer the question in your own words. A short explanation is enough.
Your feedback
Work through a hint
Hint 1
Does frozen mean not updated, or not used?
Hint 2
What does a correction need in order to change a result?
A worked explanation
The base performs the original transformations and the adapter adds its correction. Saving just the correction does not save a complete predictor.
Sources and further reading
WikiSQL dataset and evaluation rules · MLX-LM training guide · Training, validation, and test sets · Our source and experiment record