technique
Reinforcement learning with GRPO
Training the task model's weights directly against your metric: group-relative advantages, DSPy's GRPO optimizer, and why it costs far more rollouts than GEPA.
Before this
This page assumes you are comfortable with:
- techniqueBootstrapFinetuneDistilling a program into a smaller model's weights: collecting traces that pass the metric with a strong teacher, then fine-tuning the local task model on them.
- prerequisiteReinforcement learning basicsLearning from a reward instead of from correct answers: policies, rewards, advantages, and why comparing samples to each other gives a learning signal.
Why you need this
BootstrapFinetune can only teach the small model answers that a teacher already got right. When no teacher is much better, or when the metric gives partial credit that a pass or fail filter throws away, you want the model to learn from the score itself. GRPO trains the local task model's weights directly against your metric. It is stage 4's most powerful lever and by far its most expensive one.
The idea
In reinforcement learning (see Reinforcement learning basics) a policy is whatever chooses actions; here it is the language model, and an action is one full answer. The reward is a number saying how good that answer was; here it is your metric's score. There is no correct answer to copy, only scores to compare.
Group Relative Policy Optimization (GRPO) gets its learning signal by comparison inside a small group:
- Take one input. Sample answers to it from the current model, with sampling temperature above zero so they differ. is the group size.
- Score each answer with the metric, giving rewards .
- Compute the group mean and the group standard deviation (how spread out the rewards are).
- Give each answer an advantage : positive if it beat the group average, negative if it fell short, measured in standard deviations.
- Nudge the weights so answers with positive advantage become more likely and answers with negative advantage less likely, in proportion to .
The group mean plays the role of the baseline from the basics page: "better or worse than expected" where "expected" is what this model does on this input today. No separate model is needed to estimate it, which is what makes GRPO cheaper than older methods that trained a second network for the baseline.
Precisely, with population standard deviation:
Worked example
Advantages for G = 4
One math word problem, four sampled answers, scored by a metric that gives partial credit for correct working:
| Answer | Reward | Advantage | ||
|---|---|---|---|---|
| 1 | 0.90 | +0.40 | 0.16 | +1.4606 |
| 2 | 0.60 | +0.10 | 0.01 | +0.3651 |
| 3 | 0.30 | -0.20 | 0.04 | -0.7303 |
| 4 | 0.20 | -0.30 | 0.09 | -1.0954 |
Step by step: . The squared deviations sum to , so and . Then , and so on. Two checks: the advantages add to zero, and their squares add to . Answers 1 and 2 get pushed up; answers 3 and 4 get pushed down, answer 4 hardest.
Now suppose all four answers score 0.5. Then every and . Implementations usually add a tiny number to to avoid dividing by zero, so every advantage is exactly 0. A group where every answer scores the same teaches nothing. That happens when a problem is too easy (all 1.0) or too hard (all 0.0), and it is why GRPO needs a metric with room between those ends.
Move the group size, resample, then press "All equal rewards" and watch every bar collapse to zero.
DSPy's GRPO optimizer
In DSPy 3.4, GRPO is not exported as dspy.GRPO. It lives in dspy/teleprompt/grpo.py and is imported as from dspy.teleprompt.grpo import GRPO. What its source shows:
| Fact | From the DSPy 3.4 source |
|---|---|
| Group size | num_rollouts_per_grpo_step is (default 1, which gives no comparison at all). |
| Batch per step | num_dspy_examples_per_grpo_step inputs per training step, for num_train_steps steps (default 100). |
| Required settings | exclude_demos=True and multitask=True are asserted; any other value stops construction. |
| One model | The student must use a single model for all its predictors. |
| Teachers | The student must be among the teachers; if none is given, the student is its own teacher. |
| Cache | The LM cache is switched off during training, so repeated samples actually differ. |
| Unparseable output | Scored format_failure_score (default -1), below failure_score (default 0), so broken formatting is punished more than a wrong answer. |
| Who computes | DSPy sends each group of answers with their rewards to the training backend; the backend computes advantages and updates weights. |
The training backend is the catch. GRPO calls the model's reinforce method, and in DSPy 3.4 no built-in provider implements it. You need a separate training server. The DSPy reinforcement learning tutorials use Arbor (its own package, with an ArborProvider), install DSPy from its main branch, and use an ArborGRPO class whose arguments differ from 3.4's (num_samples_per_input where 3.4 says num_rollouts_per_grpo_step). The tutorial opens with "WARNING: This feature is new and extremely EXPERIMENTAL", reports running "on 4xH100 GPUs for a couple of hours" (three for training, one for serving), and its configuration sets scale_rewards to False, which skips the division by . The docs show it only on NVIDIA data-center GPUs and describe no other hardware path, so treat it as NVIDIA-only.
This file checks the 3.4 interface without training anything. Verified on DSPy 3.4.0 (ran on Python 3.14; 3.12 and newer behave the same); constructing an LM makes no network call.
import dspy
from dspy.teleprompt.grpo import GRPO
def metric(example, prediction, trace=None):
return float(example.answer == prediction.answer)
opt = GRPO(metric=metric, num_rollouts_per_grpo_step=4, num_dspy_examples_per_grpo_step=4,
num_train_steps=100, exclude_demos=True)
print("G =", opt.num_rollouts_per_grpo_step, "| rollouts planned =",
opt.num_train_steps * opt.num_dspy_examples_per_grpo_step * opt.num_rollouts_per_grpo_step)
try:
GRPO(metric=metric)
except AssertionError as err:
print("AssertionError:", err)
lm = dspy.LM("openai/local:Qwen/Qwen2.5-1.5B-Instruct") # constructing an LM makes no call
try:
lm.reinforce(train_kwargs={})
except Exception as err:
print(type(err).__name__ + ":", str(err).splitlines()[0])
Run with python construct.py:
G = 4 | rollouts planned = 1600
AssertionError: exclude_demos==False is not supported yet. Please set it to True.
LMUnsupportedFeatureError: [openai/local:Qwen/Qwen2.5-1.5B-Instruct] Provider <dspy.clients.openai.OpenAIProvider object at 0x...> does not implement the reinforcement learning interface.
A full run needs a real model and training hardware; nothing shown was run. With a reinforce-capable provider attached to student_lm, the 3.4 call is:
student = program.deepcopy()
student.set_lm(student_lm) # a model whose provider implements reinforce
opt = GRPO(metric=metric, num_rollouts_per_grpo_step=8, num_dspy_examples_per_grpo_step=4,
num_train_steps=500, exclude_demos=True, num_threads=24)
trained = opt.compile(student, trainset=trainset, valset=valset)
Multi-module programs
For a program with several predictors, every predictor runs during each rollout, and the rollout's single score becomes the reward for each predictor's part of it. With multitask=True all those groups train one shared model. If a predictor is called a different number of times in different rollouts (a loop, say), variably_invoked_predictor_grouping_mode decides whether to truncate groups to the shortest, fill them, or leave them ragged.
In an optimization pipeline
GRPO is the stage 4 lever for when the metric is the only teacher you have. Stage 2 matters even more here than for prompt optimizers, because the model is trained to raise exactly the number your metric returns. Stage 3 still comes first in practice: the GEPA paper by Agrawal and colleagues reports that GEPA beat a GRPO baseline "by up to 20% while using up to 35x fewer rollouts", where the GRPO baseline used 24,000 rollouts with LoRA on Qwen3 8B and GEPA reached its best IFBench prompt after 678 rollouts (). The DSPy GRPO tutorial itself calls its result "typically worse on cost/quality basis" than running prompt optimizers. Prompts and weights together covers running both.
Common mistakes
- Group size 1. The default
num_rollouts_per_grpo_step=1leaves nothing to compare: one answer, , advantage 0. Training runs and the model does not change. - The cache left on, or temperature 0. Every sample is identical, so every group has equal rewards. DSPy logs "GRPOGroup has no diversity".
- A metric that is all or nothing on hard inputs. Most groups score all 0.0, so most steps carry no signal. Partial credit, or easier training inputs, fixes it.
- Reward hacking. The model finds what the metric rewards rather than what you meant: a length-based metric breeds padded answers, a keyword check breeds keyword stuffing. Scores rise while a person reading the outputs sees them get worse. Read samples at every validation step.
- Formatting collapse. A small model that stops producing DSPy's field markers gets
format_failure_scoreon everything. Watch the "format failure" warnings early. - No valset. Without one, there is no honest view of whether training helps. Pass
valsetand setnum_steps_for_val.
Cost
Let be num_train_steps, the inputs per step, the group size, and the predictor calls per rollout. Training takes rollouts, so calls to the local model plus metric calls; if the metric is an LLM judge at dollars per call, that adds . The tutorial's settings (, , ) give rollouts. The rest is hardware: several GPUs for hours, at dollars per GPU-hour times GPUs times hours. Compared with BootstrapFinetune, which needs one teacher rollout per trainset example, GRPO needs samples per input on every step: a 500-example trainset costs BootstrapFinetune 500 rollouts, while the tutorial's GRPO settings cost 16,000, which is 32 times as many. Add a training server to run and monitor.
Going further
- The DSPy reinforcement learning tutorials (privacy-conscious delegation and multi-hop research), which run GRPO through Arbor.
- The GEPA paper by Agrawal and colleagues, section comparing GEPA with GRPO.
- The DeepSeekMath paper, where GRPO was introduced.
- Reinforcement learning basics, for policies, baselines, and reward hacking in general.
Leads to
Back to DSPy and GEPA: programming and optimizing language-model systems