technique

Reinforcement learning with GRPO

Training the task model's weights directly against your metric: group-relative advantages, DSPy's GRPO optimizer, and why it costs far more rollouts than GEPA.

Before this

This page assumes you are comfortable with:

Why you need this

BootstrapFinetune can only teach the small model answers that a teacher already got right. When no teacher is much better, or when the metric gives partial credit that a pass or fail filter throws away, you want the model to learn from the score itself. GRPO trains the local task model's weights directly against your metric. It is stage 4's most powerful lever and by far its most expensive one.

The idea

In reinforcement learning (see Reinforcement learning basics) a policy is whatever chooses actions; here it is the language model, and an action is one full answer. The reward rr is a number saying how good that answer was; here it is your metric's score. There is no correct answer to copy, only scores to compare.

Group Relative Policy Optimization (GRPO) gets its learning signal by comparison inside a small group:

  1. Take one input. Sample GG answers to it from the current model, with sampling temperature above zero so they differ. GG is the group size.
  2. Score each answer with the metric, giving rewards r1,…,rGr_1, \dots, r_G.
  3. Compute the group mean μ\mu and the group standard deviation σ\sigma (how spread out the rewards are).
  4. Give each answer an advantage Aj=(rj−μ)/σA_j = (r_j - \mu) / \sigma: positive if it beat the group average, negative if it fell short, measured in standard deviations.
  5. Nudge the weights so answers with positive advantage become more likely and answers with negative advantage less likely, in proportion to AjA_j.

The group mean plays the role of the baseline from the basics page: "better or worse than expected" where "expected" is what this model does on this input today. No separate model is needed to estimate it, which is what makes GRPO cheaper than older methods that trained a second network for the baseline.

Precisely, with population standard deviation:

μ=1G∑j=1Grj,σ=1G∑j=1G(rj−μ)2,Aj=rj−μσ.\mu = \frac{1}{G}\sum_{j=1}^{G} r_j, \qquad \sigma = \sqrt{\frac{1}{G}\sum_{j=1}^{G} (r_j - \mu)^2}, \qquad A_j = \frac{r_j - \mu}{\sigma}.

Worked example

Advantages for G = 4

One math word problem, four sampled answers, scored by a metric that gives partial credit for correct working:

Answer jj Reward rjr_j rj−μr_j - \mu (rj−μ)2(r_j - \mu)^2 Advantage AjA_j
1 0.90 +0.40 0.16 +1.4606
2 0.60 +0.10 0.01 +0.3651
3 0.30 -0.20 0.04 -0.7303
4 0.20 -0.30 0.09 -1.0954

Step by step: μ=(0.9+0.6+0.3+0.2)/4=2.0/4=0.5\mu = (0.9 + 0.6 + 0.3 + 0.2)/4 = 2.0/4 = 0.5. The squared deviations sum to 0.300.30, so σ2=0.30/4=0.075\sigma^2 = 0.30/4 = 0.075 and σ=0.075≈0.2739\sigma = \sqrt{0.075} \approx 0.2739. Then A1=0.40/0.2739≈1.4606A_1 = 0.40 / 0.2739 \approx 1.4606, and so on. Two checks: the advantages add to zero, and their squares add to G=4G = 4. Answers 1 and 2 get pushed up; answers 3 and 4 get pushed down, answer 4 hardest.

Now suppose all four answers score 0.5. Then every rj−μ=0r_j - \mu = 0 and σ=0\sigma = 0. Implementations usually add a tiny number to σ\sigma to avoid dividing by zero, so every advantage is exactly 0. A group where every answer scores the same teaches nothing. That happens when a problem is too easy (all 1.0) or too hard (all 0.0), and it is why GRPO needs a metric with room between those ends.

Move the group size, resample, then press "All equal rewards" and watch every bar collapse to zero.

DSPy's GRPO optimizer

In DSPy 3.4, GRPO is not exported as dspy.GRPO. It lives in dspy/teleprompt/grpo.py and is imported as from dspy.teleprompt.grpo import GRPO. What its source shows:

Fact From the DSPy 3.4 source
Group size num_rollouts_per_grpo_step is GG (default 1, which gives no comparison at all).
Batch per step num_dspy_examples_per_grpo_step inputs per training step, for num_train_steps steps (default 100).
Required settings exclude_demos=True and multitask=True are asserted; any other value stops construction.
One model The student must use a single model for all its predictors.
Teachers The student must be among the teachers; if none is given, the student is its own teacher.
Cache The LM cache is switched off during training, so repeated samples actually differ.
Unparseable output Scored format_failure_score (default -1), below failure_score (default 0), so broken formatting is punished more than a wrong answer.
Who computes AA DSPy sends each group of answers with their rewards to the training backend; the backend computes advantages and updates weights.

The training backend is the catch. GRPO calls the model's reinforce method, and in DSPy 3.4 no built-in provider implements it. You need a separate training server. The DSPy reinforcement learning tutorials use Arbor (its own package, with an ArborProvider), install DSPy from its main branch, and use an ArborGRPO class whose arguments differ from 3.4's (num_samples_per_input where 3.4 says num_rollouts_per_grpo_step). The tutorial opens with "WARNING: This feature is new and extremely EXPERIMENTAL", reports running "on 4xH100 GPUs for a couple of hours" (three for training, one for serving), and its configuration sets scale_rewards to False, which skips the division by σ\sigma. The docs show it only on NVIDIA data-center GPUs and describe no other hardware path, so treat it as NVIDIA-only.

This file checks the 3.4 interface without training anything. Verified on DSPy 3.4.0 (ran on Python 3.14; 3.12 and newer behave the same); constructing an LM makes no network call.

import dspy
from dspy.teleprompt.grpo import GRPO


def metric(example, prediction, trace=None):
    return float(example.answer == prediction.answer)


opt = GRPO(metric=metric, num_rollouts_per_grpo_step=4, num_dspy_examples_per_grpo_step=4,
           num_train_steps=100, exclude_demos=True)
print("G =", opt.num_rollouts_per_grpo_step, "| rollouts planned =",
      opt.num_train_steps * opt.num_dspy_examples_per_grpo_step * opt.num_rollouts_per_grpo_step)

try:
    GRPO(metric=metric)
except AssertionError as err:
    print("AssertionError:", err)

lm = dspy.LM("openai/local:Qwen/Qwen2.5-1.5B-Instruct")  # constructing an LM makes no call
try:
    lm.reinforce(train_kwargs={})
except Exception as err:
    print(type(err).__name__ + ":", str(err).splitlines()[0])

Run with python construct.py:

G = 4 | rollouts planned = 1600
AssertionError: exclude_demos==False is not supported yet. Please set it to True.
LMUnsupportedFeatureError: [openai/local:Qwen/Qwen2.5-1.5B-Instruct] Provider <dspy.clients.openai.OpenAIProvider object at 0x...> does not implement the reinforcement learning interface.

A full run needs a real model and training hardware; nothing shown was run. With a reinforce-capable provider attached to student_lm, the 3.4 call is:

student = program.deepcopy()
student.set_lm(student_lm)   # a model whose provider implements reinforce

opt = GRPO(metric=metric, num_rollouts_per_grpo_step=8, num_dspy_examples_per_grpo_step=4,
           num_train_steps=500, exclude_demos=True, num_threads=24)
trained = opt.compile(student, trainset=trainset, valset=valset)

Multi-module programs

For a program with several predictors, every predictor runs during each rollout, and the rollout's single score becomes the reward for each predictor's part of it. With multitask=True all those groups train one shared model. If a predictor is called a different number of times in different rollouts (a loop, say), variably_invoked_predictor_grouping_mode decides whether to truncate groups to the shortest, fill them, or leave them ragged.

In an optimization pipeline

GRPO is the stage 4 lever for when the metric is the only teacher you have. Stage 2 matters even more here than for prompt optimizers, because the model is trained to raise exactly the number your metric returns. Stage 3 still comes first in practice: the GEPA paper by Agrawal and colleagues reports that GEPA beat a GRPO baseline "by up to 20% while using up to 35x fewer rollouts", where the GRPO baseline used 24,000 rollouts with LoRA on Qwen3 8B and GEPA reached its best IFBench prompt after 678 rollouts (24000/678≈3524000 / 678 \approx 35). The DSPy GRPO tutorial itself calls its result "typically worse on cost/quality basis" than running prompt optimizers. Prompts and weights together covers running both.

Common mistakes

  • Group size 1. The default num_rollouts_per_grpo_step=1 leaves nothing to compare: one answer, σ=0\sigma = 0, advantage 0. Training runs and the model does not change.
  • The cache left on, or temperature 0. Every sample is identical, so every group has equal rewards. DSPy logs "GRPOGroup has no diversity".
  • A metric that is all or nothing on hard inputs. Most groups score all 0.0, so most steps carry no signal. Partial credit, or easier training inputs, fixes it.
  • Reward hacking. The model finds what the metric rewards rather than what you meant: a length-based metric breeds padded answers, a keyword check breeds keyword stuffing. Scores rise while a person reading the outputs sees them get worse. Read samples at every validation step.
  • Formatting collapse. A small model that stops producing DSPy's field markers gets format_failure_score on everything. Watch the "format failure" warnings early.
  • No valset. Without one, there is no honest view of whether training helps. Pass valset and set num_steps_for_val.

Cost

Let SS be num_train_steps, BB the inputs per step, GG the group size, and kk the predictor calls per rollout. Training takes R=S⋅B⋅GR = S \cdot B \cdot G rollouts, so RkR k calls to the local model plus RR metric calls; if the metric is an LLM judge at cJc_J dollars per call, that adds RcJR c_J. The tutorial's settings (S=500S = 500, B=4B = 4, G=8G = 8) give R=16000R = 16000 rollouts. The rest is hardware: several GPUs for hours, at cHc_H dollars per GPU-hour times GPUs times hours. Compared with BootstrapFinetune, which needs one teacher rollout per trainset example, GRPO needs GG samples per input on every step: a 500-example trainset costs BootstrapFinetune 500 rollouts, while the tutorial's GRPO settings cost 16,000, which is 32 times as many. Add a training server to run and monitor.

Going further

  • The DSPy reinforcement learning tutorials (privacy-conscious delegation and multi-hop research), which run GRPO through Arbor.
  • The GEPA paper by Agrawal and colleagues, section comparing GEPA with GRPO.
  • The DeepSeekMath paper, where GRPO was introduced.
  • Reinforcement learning basics, for policies, baselines, and reward hacking in general.

Leads to

Back to DSPy and GEPA: programming and optimizing language-model systems