prerequisite
Gradient descent and fine-tuning
How a model's weights change during training: a loss, its slope, a learning rate, and repeated small steps, and what fine-tuning changes compared with prompting.
Before this
This page assumes you are comfortable with:
Why you need this
Stage 4 of the pipeline changes the model itself instead of its prompt. Every method there, from BootstrapFinetune to GRPO, uses the same engine underneath: measure how wrong the model is, find which way to nudge each weight to make it less wrong, and take a small step. This page explains that engine with one weight and three steps you can check by hand.
The idea
Weights are numbers
A language model is a very large function. It takes the tokens so far and returns a probability for every possible next token. What makes it behave one way and not another is a long list of numbers called weights (or parameters): billions of them in a modern model. Training means changing those numbers. Prompting, by contrast, leaves every weight alone and changes only the text the model reads.
Loss: one number for how wrong
A loss is a single number that is small when the model does well and large when it does badly. For next-token prediction the usual loss is cross-entropy: if the model gave the correct next token probability , the loss for that token is
where is the natural logarithm. With small numbers: if the correct token got , the loss is . If it got , the loss is . A confident correct answer costs little; a hesitant one costs more; a confident wrong one (, loss ) costs most. Training on a sentence averages this over all its tokens.
Gradient: which way is downhill
(This is the one section on the page that uses calculus. If you have not met derivatives, read "slope" for "derivative" and the arithmetic still works.)
Picture the loss as a hill, with the weight's value along the ground. The gradient is the slope of the hill where you stand. For one weight it is the derivative . A positive slope means the loss rises as grows, so you should make smaller; a negative slope means the opposite. A model has a slope for every weight at once, and the list of all of them is the gradient.
Take a loss with one weight, . Its lowest point is at , where . Its derivative is
At the slope is : negative, so increasing goes downhill.
Learning rate: how big a step
Gradient descent repeats one rule: move each weight a little against its slope.
The learning rate (Greek letter eta) sets the step size. Too small and training crawls. Too large and each step overshoots the bottom and lands higher on the other side.
Epochs and batches
Real training computes the loss on a batch of examples at a time (say 8), takes one step, then moves to the next batch. One pass through the whole training set is an epoch. With 1,000 examples and batches of 8, one epoch is steps, and 5 epochs are 625 steps.
Worked example
Three steps of gradient descent on , starting at with :
| Step | before | Slope | Update slope | after | after |
|---|---|---|---|---|---|
| start | 0 | 9 | |||
| 1 | 0 | -6 | +0.6 | 0.6 | 5.76 |
| 2 | 0.6 | -4.8 | +0.48 | 1.08 | 3.6864 |
| 3 | 1.08 | -3.84 | +0.384 | 1.464 | 2.359296 |
Each step is smaller than the last, because the slope shrinks as you approach the bottom. The same steps in Python (written for this page, runs with no model):
"""Three gradient-descent steps on L(w) = (w - 3)**2."""
def loss(w):
return (w - 3) ** 2
def slope(w):
return 2 * (w - 3)
w, eta = 0.0, 0.1
print(f"start w={w:.4f} L={loss(w):.4f}")
for step in range(1, 4):
g = slope(w)
w = w - eta * g
print(f"step {step} slope={g:+.4f} w={w:.4f} L={loss(w):.4f}")
Run with python steps.py (Python 3.14; 3.12 and newer behave the same):
start w=0.0000 L=9.0000
step 1 slope=-6.0000 w=0.6000 L=5.7600
step 2 slope=-4.8000 w=1.0800 L=3.6864
step 3 slope=-3.8400 w=1.4640 L=2.3593
Now change the learning rate. With the first step jumps from 0 to 6.6, past the bottom at 3, and the loss rises from 9 to 12.96; the next steps go to -1.32 and 8.184, with losses 18.66 and 26.87. The weight bounces further out each time. With three steps only reach , with the loss still at 7.97. Real training picks by trying a few values; fine-tuning a large model uses very small rates, such as the 0.00001 that DSPy 3.4's local fine-tuning provider uses by default.
Supervised fine-tuning
Fine-tuning starts from a model that is already trained and continues training it on your own data. Supervised fine-tuning (SFT) uses input and output pairs: a ticket and its correct label, a question and its worked answer. The loss is cross-entropy on the output tokens only, so the model is pushed to give high probability to your answers when it sees your inputs.
| Prompting | Fine-tuning | |
|---|---|---|
| What changes | The text the model reads | The weights |
| Needs | A few examples | Hundreds or more pairs |
| Takes effect | Immediately | After a training run on a GPU |
| Undo | Edit the prompt | Keep the old weights |
In DSPy the pairs come from traces: BootstrapFinetune runs a strong teacher, keeps runs that pass the metric, and trains the student on them. The model being trained is a checkpoint you hold the weights for, loaded by a training library; DSPy 3.4 does not fine-tune a model served by Ollama.
Overfitting and catastrophic forgetting
Two things go wrong. Overfitting: after too many epochs on a small set, the model memorizes those exact examples and does worse on new ones; the training loss keeps falling while the validation loss starts rising. Catastrophic forgetting: training hard on one narrow task can erode abilities the model had before, such as following other instructions or writing in a normal tone. Both are why fine-tuning uses small learning rates, few epochs, a held-out valset, and sometimes trains only a small add-on instead of every weight, which is the subject of LoRA and parameter-efficient tuning.
In an optimization pipeline
Stage 3 (prompt optimizers) never computes a gradient: it searches over text. Stage 4 does. BootstrapFinetune is supervised fine-tuning on filtered traces, and Reinforcement learning basics explains how a reward replaces the correct answer when you have only a score. Both end in gradient steps like the three above, repeated millions of times over billions of weights.
Common mistakes
- Learning rate too high. The loss jumps around or grows, like the run above. Symptom: training loss that does not fall, or outputs that turn to gibberish.
- Too many epochs on a small set. Training loss near zero, validation score dropping. The model repeats training answers word for word.
- Judging by training loss. A low training loss says the model fits the training pairs, not that it is useful. Score the tuned program with your metric on a valset.
- Forgetting the old skills. A model tuned only on ticket labels may stop answering anything else well. Keep a few general checks in your evaluation.
Cost
Training costs memory and time. For each weight, training holds the weight, its gradient, and usually two more numbers kept by the optimizer, so memory is several times what serving the model needs. Time grows with the number of weights times the number of training tokens times the number of epochs. In practice this means a GPU, and the larger the model, the larger the GPU; parameter-efficient methods exist to cut that cost.
Going further
- LoRA and parameter-efficient tuning, for training a small add-on instead of every weight.
- Reinforcement learning basics, for learning from a score instead of a correct answer.
- BootstrapFinetune, for DSPy's fine-tuning optimizer.
- The idea of the Adam optimizer, the gradient-descent variant most fine-tuning uses, which keeps a running average of each weight's slope.
Leads to
- prerequisiteLoRA and parameter-efficient tuningFine-tuning a large model by training a small add-on instead of every weight: low-rank adapters, their size, and what they cost in memory.
- prerequisiteReinforcement learning basicsLearning from a reward instead of from correct answers: policies, rewards, advantages, and why comparing samples to each other gives a learning signal.
Back to DSPy and GEPA: programming and optimizing language-model systems