prerequisite

Gradient descent and fine-tuning

How a model's weights change during training: a loss, its slope, a learning rate, and repeated small steps, and what fine-tuning changes compared with prompting.

Before this

This page assumes you are comfortable with:

Why you need this

Stage 4 of the pipeline changes the model itself instead of its prompt. Every method there, from BootstrapFinetune to GRPO, uses the same engine underneath: measure how wrong the model is, find which way to nudge each weight to make it less wrong, and take a small step. This page explains that engine with one weight and three steps you can check by hand.

The idea

Weights are numbers

A language model is a very large function. It takes the tokens so far and returns a probability for every possible next token. What makes it behave one way and not another is a long list of numbers called weights (or parameters): billions of them in a modern model. Training means changing those numbers. Prompting, by contrast, leaves every weight alone and changes only the text the model reads.

Loss: one number for how wrong

A loss LL is a single number that is small when the model does well and large when it does badly. For next-token prediction the usual loss is cross-entropy: if the model gave the correct next token probability pp, the loss for that token is

L=−log⁡pL = -\log p

where log⁡\log is the natural logarithm. With small numbers: if the correct token got p=0.25p = 0.25, the loss is −log⁡0.25≈1.386-\log 0.25 \approx 1.386. If it got p=0.9p = 0.9, the loss is −log⁡0.9≈0.105-\log 0.9 \approx 0.105. A confident correct answer costs little; a hesitant one costs more; a confident wrong one (p=0.01p = 0.01, loss ≈4.605\approx 4.605) costs most. Training on a sentence averages this over all its tokens.

Gradient: which way is downhill

(This is the one section on the page that uses calculus. If you have not met derivatives, read "slope" for "derivative" and the arithmetic still works.)

Picture the loss as a hill, with the weight's value along the ground. The gradient is the slope of the hill where you stand. For one weight ww it is the derivative dLdw\frac{dL}{dw}. A positive slope means the loss rises as ww grows, so you should make ww smaller; a negative slope means the opposite. A model has a slope for every weight at once, and the list of all of them is the gradient.

Take a loss with one weight, L(w)=(w−3)2L(w) = (w - 3)^2. Its lowest point is at w=3w = 3, where L=0L = 0. Its derivative is

dLdw=2(w−3)\frac{dL}{dw} = 2(w - 3)

At w=0w = 0 the slope is 2(0−3)=−62(0 - 3) = -6: negative, so increasing ww goes downhill.

Learning rate: how big a step

Gradient descent repeats one rule: move each weight a little against its slope.

wnew=w−η⋅dLdww_{\text{new}} = w - \eta \cdot \frac{dL}{dw}

The learning rate η\eta (Greek letter eta) sets the step size. Too small and training crawls. Too large and each step overshoots the bottom and lands higher on the other side.

Epochs and batches

Real training computes the loss on a batch of examples at a time (say 8), takes one step, then moves to the next batch. One pass through the whole training set is an epoch. With 1,000 examples and batches of 8, one epoch is 1000/8=1251000 / 8 = 125 steps, and 5 epochs are 625 steps.

Worked example

Three steps of gradient descent on L(w)=(w−3)2L(w) = (w - 3)^2, starting at w=0w = 0 with η=0.1\eta = 0.1:

Step ww before Slope 2(w−3)2(w-3) Update −η⋅-\eta \cdot slope ww after LL after
start 0 9
1 0 -6 +0.6 0.6 5.76
2 0.6 -4.8 +0.48 1.08 3.6864
3 1.08 -3.84 +0.384 1.464 2.359296

Each step is smaller than the last, because the slope shrinks as you approach the bottom. The same steps in Python (written for this page, runs with no model):

"""Three gradient-descent steps on L(w) = (w - 3)**2."""


def loss(w):
    return (w - 3) ** 2


def slope(w):
    return 2 * (w - 3)


w, eta = 0.0, 0.1
print(f"start   w={w:.4f}  L={loss(w):.4f}")
for step in range(1, 4):
    g = slope(w)
    w = w - eta * g
    print(f"step {step}  slope={g:+.4f}  w={w:.4f}  L={loss(w):.4f}")

Run with python steps.py (Python 3.14; 3.12 and newer behave the same):

start   w=0.0000  L=9.0000
step 1  slope=-6.0000  w=0.6000  L=5.7600
step 2  slope=-4.8000  w=1.0800  L=3.6864
step 3  slope=-3.8400  w=1.4640  L=2.3593

Now change the learning rate. With η=1.1\eta = 1.1 the first step jumps from 0 to 6.6, past the bottom at 3, and the loss rises from 9 to 12.96; the next steps go to -1.32 and 8.184, with losses 18.66 and 26.87. The weight bounces further out each time. With η=0.01\eta = 0.01 three steps only reach w=0.176w = 0.176, with the loss still at 7.97. Real training picks η\eta by trying a few values; fine-tuning a large model uses very small rates, such as the 0.00001 that DSPy 3.4's local fine-tuning provider uses by default.

Supervised fine-tuning

Fine-tuning starts from a model that is already trained and continues training it on your own data. Supervised fine-tuning (SFT) uses input and output pairs: a ticket and its correct label, a question and its worked answer. The loss is cross-entropy on the output tokens only, so the model is pushed to give high probability to your answers when it sees your inputs.

Prompting Fine-tuning
What changes The text the model reads The weights
Needs A few examples Hundreds or more pairs
Takes effect Immediately After a training run on a GPU
Undo Edit the prompt Keep the old weights

In DSPy the pairs come from traces: BootstrapFinetune runs a strong teacher, keeps runs that pass the metric, and trains the student on them. The model being trained is a checkpoint you hold the weights for, loaded by a training library; DSPy 3.4 does not fine-tune a model served by Ollama.

Overfitting and catastrophic forgetting

Two things go wrong. Overfitting: after too many epochs on a small set, the model memorizes those exact examples and does worse on new ones; the training loss keeps falling while the validation loss starts rising. Catastrophic forgetting: training hard on one narrow task can erode abilities the model had before, such as following other instructions or writing in a normal tone. Both are why fine-tuning uses small learning rates, few epochs, a held-out valset, and sometimes trains only a small add-on instead of every weight, which is the subject of LoRA and parameter-efficient tuning.

In an optimization pipeline

Stage 3 (prompt optimizers) never computes a gradient: it searches over text. Stage 4 does. BootstrapFinetune is supervised fine-tuning on filtered traces, and Reinforcement learning basics explains how a reward replaces the correct answer when you have only a score. Both end in gradient steps like the three above, repeated millions of times over billions of weights.

Common mistakes

  • Learning rate too high. The loss jumps around or grows, like the η=1.1\eta = 1.1 run above. Symptom: training loss that does not fall, or outputs that turn to gibberish.
  • Too many epochs on a small set. Training loss near zero, validation score dropping. The model repeats training answers word for word.
  • Judging by training loss. A low training loss says the model fits the training pairs, not that it is useful. Score the tuned program with your metric on a valset.
  • Forgetting the old skills. A model tuned only on ticket labels may stop answering anything else well. Keep a few general checks in your evaluation.

Cost

Training costs memory and time. For each weight, training holds the weight, its gradient, and usually two more numbers kept by the optimizer, so memory is several times what serving the model needs. Time grows with the number of weights times the number of training tokens times the number of epochs. In practice this means a GPU, and the larger the model, the larger the GPU; parameter-efficient methods exist to cut that cost.

Going further

Leads to

Back to DSPy and GEPA: programming and optimizing language-model systems