prerequisite

Reinforcement learning basics

Learning from a reward instead of from correct answers: policies, rewards, advantages, and why comparing samples to each other gives a learning signal.

Before this

This page assumes you are comfortable with:

Why you need this

Supervised fine-tuning needs the correct output for every input. Often you do not have it: you have a metric that can score an answer, but nobody wrote the ideal answer. Reinforcement learning trains the model from those scores alone. It is the idea behind GRPO in stage 4 of the pipeline, and its central trick, comparing an answer with the model's other answers, explains why GRPO needs many samples per input.

The idea

The words

Term Meaning here
Policy The model, seen as a rule for choosing what to say. For a language model it is the probability distribution over outputs for a given prompt.
Action One output the policy produced: a sampled answer.
Reward rr A number scoring that action. In DSPy it is your metric, a value in [0, 1].
Baseline bb What reward the policy usually gets on this prompt: the expected reward.
Advantage AA How much better or worse this action did than expected: A=r−bA = r - b.

Supervised versus reinforcement

In supervised fine-tuning every training pair says "for this input, say exactly this." The loss pulls the model toward that one answer. In reinforcement learning no one says what to answer. The model samples its own answers, a reward function scores them, and training makes high-scoring answers more likely and low-scoring ones less likely. The model can only learn from answers it actually produces, so it must sometimes produce good ones by chance.

Better or worse than expected

A raw reward is a poor training signal by itself. If every answer to a prompt scores between 0.8 and 1.0, pushing up every answer in proportion to its reward pushes up even the worst one. What matters is whether an answer did better or worse than the policy normally does. That is the advantage A=r−bA = r - b:

  • A>0A > 0: better than expected. Make this answer more likely.
  • A<0A < 0: worse than expected. Make it less likely.
  • A=0A = 0: exactly as expected. Leave it alone.

A simple and common baseline is the average reward of several answers sampled for the same prompt. Then the answers are graded against each other.

The policy gradient, without new calculus

The previous page moved a weight against the slope of a loss. Reinforcement learning uses the same gradient step with a different target: for each sampled answer, nudge the weights so that answer's log-probability changes by an amount proportional to its advantage. In symbols, the step on answer ii is about η⋅Ai\eta \cdot A_i times the direction that makes answer ii more likely, where η\eta is the learning rate. Positive advantage, that answer's probability goes up; negative, it goes down; the bigger ∣Ai∣|A_i|, the bigger the push. Averaged over many prompts and samples, the model drifts toward whatever the reward prefers.

Worked example

Prompt: a support ticket, "I was charged twice this month. Reply with one label: billing, bug, or account." The policy samples four answers. The metric gives 1 for exactly the right label, 0.5 for the right label wrapped in extra words, and 0 for a wrong label.

Answer ii Sampled output Reward rir_i
1 billing 1
2 bug 0
3 billing issue, refund needed 0.5
4 billing 1

Baseline, the average reward: b=(1+0+0.5+1)/4=2.5/4=0.625b = (1 + 0 + 0.5 + 1) / 4 = 2.5 / 4 = 0.625.

Answer rir_i Ai=ri−bA_i = r_i - b Update
1 1 +0.375 more likely
2 0 -0.625 less likely, the biggest push
3 0.5 -0.125 slightly less likely
4 1 +0.375 more likely

Notice answer 3. Its reward is positive, so using raw rewards would have made it more likely. Against the baseline it is below average, so the model is pushed away from wrapping the label in extra words. Notice also what happens if all four answers had scored 1: the baseline would be 1, every advantage 0, and nothing would be learned from this prompt. Groups where every answer gets the same reward carry no signal, which is why the sample count per prompt matters.

In an optimization pipeline

In stage 4, Reinforcement learning with GRPO applies exactly this: sample a group of GG outputs per input, use the metric as the reward, and compute each output's advantage against the group, with one extra step (dividing by the group's spread) that the GRPO page explains. The metric you wrote in stage 2 becomes the reward, so every weakness in it becomes something the model can learn to exploit.

Reward hacking

Reward hacking is the policy finding answers that score well without doing the task. Suppose the triage metric gave reward 1 whenever the output contains the right label, instead of equals it. Then the answer "billing bug account" contains every label and scores 1 on every ticket. Sampled once by accident, it gets a positive advantage, becomes more likely, and soon the model answers every ticket with all three labels. The training curve shows the reward climbing to 1; the outputs are useless. A prompt optimizer can exploit a loose metric too, but reinforcement learning changes the weights, and it does so over thousands of steps, so it is far more persistent at finding the hole.

Common mistakes

  • No baseline. Pushing every answer up by its raw reward makes training slow and noisy, because mediocre answers are rewarded too. Symptom: the average reward barely rises.
  • A metric that can be gamed. Symptom: the reward reaches near 1 while a person reading the outputs sees nonsense or a repeated trick.
  • All-equal rewards. Prompts that are always solved, or never solved, give zero advantage. Symptom: a training run that spends compute and does not move. Use inputs the model gets right some of the time.
  • Too few samples per prompt. With two samples the baseline is a noisy guess. Symptom: the reward curve jumps around from step to step.

Cost

Let PP be the number of training prompts, GG the samples per prompt, and EE the number of passes over the prompts. Training generates P⋅G⋅EP \cdot G \cdot E answers, scores each with the metric, and takes a gradient step on them. Generation dominates: every answer is a full run of the model, and a metric that itself calls a model (an LLM judge) adds a call per answer. Compared with supervised fine-tuning on the same prompts, that is roughly GG times as many model runs, plus everything that makes training expensive on the previous page: a GPU large enough to hold the model, its gradients, and optimizer state.

Going further

  • Reinforcement learning with GRPO, for group-relative advantages and DSPy's GRPO optimizer.
  • Writing metrics, for metrics that are hard to game.
  • LoRA and parameter-efficient tuning, for keeping the training cost down.
  • The GEPA paper by Agrawal and colleagues, which compares reflective prompt evolution with GRPO and discusses why a plain-language reason can teach more per rollout than a single reward number.

Leads to

Back to DSPy and GEPA: programming and optimizing language-model systems