technique

GEPA: reflective prompt evolution

How GEPA improves prompts by reading what went wrong: run a minibatch, reflect on traces and feedback in plain language, propose a new instruction, and keep a Pareto front of candidates.

Before this

This page assumes you are comfortable with:

Why you need this

You have a program that works most of the time, a metric, and a valset. Hand-editing the instruction is the slow loop from the hub: fix one case, break another. GEPA automates that loop the way a careful person would do it. It runs the program on a few examples, reads why the outputs were wrong, and rewrites the instruction to fix those reasons, keeping every version that is best at something. It is the main stage 3 optimizer in this cluster. This page explains the algorithm; GEPA in practice covers the dspy.GEPA settings.

The idea

In plain words: GEPA is an evolutionary algorithm (see Evolutionary algorithms) whose genomes are instructions and whose mutation is done by a strong language model that reads written feedback. The name stands for Genetic-Pareto. It was introduced in the GEPA paper by Agrawal and colleagues (2025), "GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning".

Some terms first. A program is a DSPy module tree; a predictor is one model call inside it, with its own instruction. A candidate is one version of the program, written as one instruction per predictor. A rollout is one run of the program on one example, and a trace is the record of every predictor's inputs and outputs during that rollout. GEPA uses two models: the local task model (a small model served by Ollama) runs the program, and the reflection model (Claude Opus 5.5) writes new instructions.

The paper splits the training data into two sets: DfeedbackD_{\text{feedback}}, from which small minibatches are drawn for reflection, and DparetoD_{\text{pareto}}, which the paper calls "the validation set used for selection". In DSPy these are your trainset and valset. The candidate pool starts with one member, your program as written. Then each iteration does this:

  1. Pick a parent from the pool using per-example Pareto selection (the paper's Section 3.1). For each valset example, find the candidates with the best score; drop candidates that are dominated; sample one, weighted by how many examples it leads. Pareto fronts works this through.
  2. Pick a predictor to change. The paper and DSPy's default use round-robin: each candidate cycles through its predictors in order, one per iteration.
  3. Run a minibatch. Draw bb examples from the trainset (b=3b = 3 by default in DSPy 3.4, the reflection_minibatch_size argument) and run the parent on them, capturing traces. The metric returns, for each rollout, a score and a feedback text that says what was wrong. The paper calls this metric-with-feedback μf\mu_f; its example of feedback is compiler error messages.
  4. Reflect. The reflection model is shown, in the paper's words, the "current prompt, language program trajectory, score, feedback" and asked for a better instruction for the chosen predictor. This is reflective mutation. If every minibatch score is already perfect, DSPy skips this step by default (skip_perfect_score=True), since there is nothing to learn from.
  5. Re-score the child on the same bb examples. In the installed gepa 0.1.4 source, the default acceptance rule is that the sum of the child's minibatch scores must be strictly greater than the parent's. Otherwise the child is thrown away.
  6. Evaluate an accepted child on the whole valset, add it to the pool, and update each example's best score. This full evaluation is what makes the Pareto selection in step 1 possible.

Sometimes an iteration does a merge instead, the paper's "system-aware" crossover. It only helps a program with several predictors. In the gepa 0.1.4 source, merge takes two candidates from the Pareto set that share an ancestor, and for each predictor uses the instruction from whichever of the two changed it relative to that ancestor. If one line of descent improved the "extract facts" predictor and another the "answer" predictor, the child gets both. It is tested on a few valset examples (5 by default in the source) and kept if it scores at least as well as the better parent. DSPy turns merge on by default (use_merge=True) and allows 5 merges per run by default (max_merge_invocations=5); a merge is only scheduled after an iteration that added a new candidate.

Why read feedback instead of a score? The paper's argument is that "the interpretable nature of language often provides a much richer learning medium for LLMs, compared to policy gradients derived from sparse, scalar rewards." Reinforcement learning such as GRPO gets one number per rollout and needs many rollouts to learn what it rewards. A trace plus "Expected 3, got 3.33: round down for full boxes" says what to change after a single failure. The paper's abstract (revised version) reports that across six tasks GEPA outperforms GRPO by 6% on average and by up to 20% while using up to 35 times fewer rollouts, and that it outperforms MIPROv2 by over 10%, for example +12% accuracy on AIME-2025. Its Section 4 also reports that GEPA's prompts are up to 9.2 times shorter than MIPROv2's, that GEPA with merge can beat GEPA without it by as much as 5%, and that Pareto-based parent selection did better than always picking the best candidate, which the paper says "often traps the optimizer in a local optimum". These are the paper's results, not guarantees for your task.

Worked example

The task is short math word problems that need a bare numeric answer. The program has one predictor whose instruction is "Solve the math word problem." This file runs one GEPA iteration by hand on a three-example minibatch. The task model is DSPy's fake model, DummyLM, scripted with the answers a weak model might give, and the reflection model is replaced by a plain function returning a scripted proposal. Everything else is real: the feedback metric has the shape dspy.GEPA expects, the reflection prompt is rendered by gepa's own InstructionProposalSignature (the class DSPy 3.4 uses to propose instructions), and the accept decision comes from gepa's default acceptance rule. It ran on DSPy 3.4 with gepa 0.1.4 and Python 3.14 (3.12 and newer behave the same), and no network call was made. The reflection prompt uses triple backticks as fences; the script prints them as ''' so they display cleanly here.

import dspy
from dspy.utils import DummyLM
from gepa.proposer.base import CandidateProposal
from gepa.strategies.acceptance import StrictImprovementAcceptance
from gepa.strategies.instruction_proposal import InstructionProposalSignature

minibatch = [
    dspy.Example(question="A pen costs $1.25. How much do 10 pens cost?", answer="12.50"),
    dspy.Example(question="Eggs come in boxes of 12. How many full boxes can you fill with 40 eggs?", answer="3"),
    dspy.Example(question="A train goes 60 km per hour for 2.5 hours. How many km does it travel?", answer="150"),
]
minibatch = [ex.with_inputs("question") for ex in minibatch]

parent_text = "Solve the math word problem."
program = dspy.Predict(dspy.Signature("question -> answer", parent_text))

def metric(gold, pred, trace=None, pred_name=None, pred_trace=None):
    got = pred.answer.strip()
    if got == gold.answer:
        return dspy.Prediction(score=1.0, feedback="Correct.")
    if gold.answer == "12.50":
        why = "The arithmetic is right but the format is wrong: give a bare number with no currency sign, and write money with two decimal places."
    elif gold.answer == "3":
        why = "The question asks for full boxes, so round down to a whole number."
    else:
        why = "Recheck the arithmetic."
    return dspy.Prediction(score=0.0, feedback=f"Expected {gold.answer}, got {got!r}. {why}")

def run(prog, answers):
    """Run prog on the minibatch with the task model scripted to give `answers`."""
    lm = DummyLM({ex.question: {"answer": a} for ex, a in zip(minibatch, answers)})
    with dspy.context(lm=lm):
        preds = [prog(question=ex.question) for ex in minibatch]
    return preds, [metric(ex, p) for ex, p in zip(minibatch, preds)]

# 1. Rollouts of the parent on the minibatch: outputs, scores, feedback.
preds, results = run(program, ["$12.5", "3.33", "150"])
for ex, p, r in zip(minibatch, preds, results):
    print(f"{p.answer!r:>8}  score {r.score}  {r.feedback}")
before = [r.score for r in results]

# 2. The reflective dataset, in the shape dspy.GEPA builds for one predictor.
records = [{"Inputs": {"question": ex.question}, "Generated Outputs": {"answer": p.answer}, "Feedback": r.feedback}
           for ex, p, r in zip(minibatch, preds, results)]

# 3. Reflection: GEPA's own prompt template, answered by a scripted stand-in for the reflection model.
FENCE = "`" * 3   # three backticks: the reflection model must wrap its answer in them

def scripted_reflection_model(prompt):
    return (FENCE + "\nSolve the math word problem. Reply with a bare number only: no units and no currency signs. "
            "Write money amounts with exactly two decimal places (12.50). When the question asks how many full "
            "or complete groups fit, round down to a whole number.\n" + FENCE)

out, prompt, _ = InstructionProposalSignature.run_with_metadata(
    lm=scripted_reflection_model,
    input_dict={"current_instruction_doc": parent_text, "dataset_with_feedback": records},
)
child_text = out["new_instruction"]
print("\n--- reflection prompt (backtick fences printed as ''') ---\n" + prompt.replace(FENCE, "'''"))
print("\n--- proposed instruction ---\n" + child_text)

# 4. Rollouts of the child on the SAME minibatch.
child = dspy.Predict(dspy.Signature("question -> answer", child_text))
_, child_results = run(child, ["12.50", "3", "150"])
after = [r.score for r in child_results]

# 5. GEPA's default acceptance rule: the minibatch sum must strictly improve.
proposal = CandidateProposal(candidate={"predict": child_text}, parent_program_ids=[0],
                             subsample_scores_before=before, subsample_scores_after=after)
print(f"\nminibatch before {sum(before)}, after {sum(after)}, accepted: "
      f"{StrictImprovementAcceptance().should_accept(proposal, state=None)}")

Run it with python one_iteration.py. It printed (the middle of the reflection prompt is trimmed at the ...):

 '$12.5'  score 0.0  Expected 12.50, got '$12.5'. The arithmetic is right but the format is wrong: give a bare number with no currency sign, and write money with two decimal places.
  '3.33'  score 0.0  Expected 3, got '3.33'. The question asks for full boxes, so round down to a whole number.
   '150'  score 1.0  Correct.

--- reflection prompt (backtick fences printed as ''') ---
I provided an assistant with the following instructions to perform a task for me:
'''
Solve the math word problem.
'''

The following are examples of different task inputs provided to the assistant along with the assistant's response for each of them, and some feedback on how the assistant's response could be better:
'''
# Example 1
## Inputs
### question
A pen costs $1.25. How much do 10 pens cost?

## Generated Outputs
### answer
$12.5

## Feedback
Expected 12.50, got '$12.5'. The arithmetic is right but the format is wrong: give a bare number with no currency sign, and write money with two decimal places.

...

## Feedback
Correct.


'''

Your task is to write a new instruction for the assistant.

Read the inputs carefully and identify the input format and infer detailed task description about the task I wish to solve with the assistant.

Read all the assistant responses and the corresponding feedback. Identify all niche and domain specific factual information about the task and include it in the instruction, as a lot of it may not be available to the assistant in the future. The assistant may have utilized a generalizable strategy to solve the task, if so, include that in the instruction as well.

Provide the new instructions within ''' blocks.

--- proposed instruction ---
Solve the math word problem. Reply with a bare number only: no units and no currency signs. Write money amounts with exactly two decimal places (12.50). When the question asks how many full or complete groups fit, round down to a whole number.

minibatch before 1.0, after 3.0, accepted: True

Step by step, this is one iteration:

Step What happened
Instruction before "Solve the math word problem."
Minibatch 1 of 3 right. The pen answer has the right value in the wrong format; the eggs answer did not round down.
Feedback Names both missing rules in plain words.
Reflection prompt The template above: the current instruction, then each example's inputs, outputs, and feedback, then a request to put task-specific facts into the new instruction.
Instruction after States the bare-number, two-decimals, and round-down rules.
Re-score 3 of 3 on the same minibatch. 3>13 > 1, so the child is accepted.
Next The child is scored on the full valset, say 62% against the parent's 50% on a 50-example valset (an illustration; nothing here ran on a valset), and joins the pool.

The metric's feedback already said what was wrong; the reflection model's job was to turn specific complaints into general rules. With score-only feedback ("This trajectory got a score of 0.0.", which DSPy 3.4 substitutes when a metric returns a bare number) the reflection model can see that the pen answer failed but must guess why.

The demo below steps through the same loop on a different task, sorting support tickets into billing, bug, account, or other. Its "score-only feedback" toggle shows what happens to the proposal when the feedback is just numbers.

In an optimization pipeline

GEPA sits in stage 3 and leans on stage 2 more than any other optimizer. The metric must return feedback, not just a score (see Writing metrics and LLM-as-judge metrics). The trainset feeds minibatches, so make it large and varied. If you pass no valset, DSPy reuses the trainset as the valset, which overfits. The valset decides which candidates survive as parents and should cover every kind of input you care about, because each example is an objective. A program built from named predictors, as in Composing programs, lets GEPA target and merge each predictor separately. After the run, the candidate with the best valset average is the one you keep, and its testset score is the one you report.

Common mistakes

  • Score-only feedback. Symptom: proposals are generic ("read carefully", "be accurate") and most are rejected, so the budget is spent without progress.
  • Feedback that leaks the answer. "Expected 12.50" is fine; feedback that pastes the whole worked solution teaches the reflection model to memorize valset items. Symptom: big valset gains, flat testset.
  • A valset too small or too narrow. Symptom: the Pareto set holds one or two candidates and the search behaves like a single hill climber.
  • A weak reflection model. Symptom: proposals repeat the old instruction with small rewording, or drop rules that were working.
  • Trusting the minibatch. Three examples is a tiny test; an accepted child can still be worse on the valset. Symptom: many accepted children that never become the best candidate.

Cost

Count metric calls, because GEPA's budget is a number of them. With minibatch size bb and a valset of nvn_v examples, one reflective iteration costs bb rollouts for the parent, bb for the child, and, only if the child is accepted, nvn_v more for the full evaluation. With b=3b = 3 and nv=50n_v = 50, a rejected iteration costs 6 metric calls and an accepted one 56. Over II iterations of which AA are accepted, the total is about 2bI+nvA2bI + n_v A metric calls, plus the starting program's nvn_v. Each metric call is one program run (one task-model call per predictor) plus a judge call if the metric uses one. On top of that, each iteration makes one reflection-model call whose input holds the instruction and bb full traces. In dollars, with reflection-model prices cinc_{\text{in}} and coutc_{\text{out}} per token and tint_{\text{in}}, toutt_{\text{out}} tokens per reflection call, reflection costs about I(tincin+toutcout)I (t_{\text{in}} c_{\text{in}} + t_{\text{out}} c_{\text{out}}), and the task model runs locally. No GPU training is involved; the usual bottlenecks are task-model wall-clock time and the builder's time writing good feedback.

Going further

  • GEPA in practice: budgets, the reflection model setting, and reading the results.
  • MIPROv2, the optimizer the paper compares against for prompts, and Reinforcement learning with GRPO, the one it compares against for weights.
  • The GEPA paper, Section 3 (the algorithm) and Section 4 (its observations).
  • The standalone gepa library, which applies the same loop to any text you can score, not only DSPy instructions.

Leads to

Back to DSPy and GEPA: programming and optimizing language-model systems