Lesson 1 of 5 · Fine-tuning for SQL

Ask a pretrained model for SQL

A model can write valid SQL and still answer the wrong question. Before changing its weights, find out what actually fails.

A useful task for a model that already knows code

In the Shakespeare course, we trained a small transformer from random numbers. Here we start with Qwen2.5-Coder-0.5B-Instruct, a pretrained, instruction-tuned model with about 494 million parameters. It already has language and code skills. Our question is whether examples can help it follow one particular task.

Supervised fine-tuning (SFT) continues training on examples of an input and a desired response. Our input is a question plus a table schema—the column names and types. Our desired response is one SQL query. We use WikiSQL, an existing dataset of questions, tables, and reference queries.

This lab supports one table, one selected column, optional aggregation such as SUM or MIN, and simple filters joined by AND. Joins and arbitrary database operations are outside the task.

Read the table before judging the answer

Our real development example asks: “How many laps did Ricardo Zonta have?” Below is the matching row from WikiSQL table 2-1123405-2. The learner can inspect it here; the model received only the schema and question.

One matching source row, displayed as stored in the database
ColumnStored value
col0 · Driverricardo zonta
col1 · Constructorbar - honda
col2 · Laps53
col3 · Time/Retired+1:09.293
col4 · Grid17

SELECT col2 retrieves the Laps field. WHERE col0 = 'ricardo zonta' selects the matching driver. COUNT(*) counts matching rows; it does not read the number of laps. Lowercase text and identifiers such as col0 are conventions of this dataset, not requirements for SQL in general.

Compare the saved queries

Recorded experiment · These controls inspect saved outputs or illustrate the procedure. They do not run a model or train in your browser.

SELECT COUNT(*) FROM data WHERE col0 = 'Ricardo Zonta'

Recorded answer: 0. The original query uses a capitalized name that does not match the stored text. Even with lowercase text, COUNT(*) would return 1 row, not 53 laps.

The adapter returns 53, but the reference uses SUM(col2). Retrieving a value and summing values agree for this one matching row; they can disagree if more rows match. One correct executed answer does not establish that the queries are equivalent.

Reserve the checks before experimenting

We selected 2,000 examples from the official training split to supply updates, 100 from the development split to guide choices, and 200 from the test split for the final comparison. The source splits have different table IDs. Training and checking different questions on the same table would be a weaker check of transfer to new tables.

Run the original model first. Try a reasonable prompt with examples drawn from training data, and inspect mistakes on development data. Fixing an input convention may be cheaper than fine-tuning. Decide the scoring rules and comparison budget before looking at test scores.

This course replays a completed experiment. Its published test questions are now historical evidence. Use development data for further changes; a new claim of transfer needs a fresh final check.

Your first practical decision

Classify a failure before choosing a fix: did the model use nonexistent columns, the wrong value, the wrong operation, or the wrong output format? A loss curve cannot tell these apart. In our example both value matching and the requested operation matter.

Download the source and recorded results. The browser path needs no installation or account. The optional local lab uses Apple Silicon and Python 3.12; setup and guarded commands appear in lesson 4. Inspect the complete original development outputs.

Your turn

Explain it in your own words.

Not checked

One matching row has Driver = ricardo zonta and Laps = 53. A query with the correct lowercase filter returns COUNT(*) = 1. Does it answer “How many laps did Ricardo Zonta have?” Explain what the query measured and what should be read instead.

Answer the question in your own words. A short explanation is enough.

Draft saves on this device0/800 characters

Work through a hint

Sources and further reading

WikiSQL dataset and evaluation rules · MLX-LM training guide · Training, validation, and test sets · Our source and experiment record