technique
GEPA in practice
Running dspy.GEPA: the metric with feedback, the reflection model, budgets, multi-module programs, reading the results, and the standalone gepa library for optimizing any text.
Before this
This page assumes you are comfortable with:
- techniqueGEPA: reflective prompt evolutionHow GEPA improves prompts by reading what went wrong: run a minibatch, reflect on traces and feedback in plain language, propose a new instruction, and keep a Pareto front of candidates.
- techniqueConfiguring language modelsConnecting DSPy to a local task model served by Ollama and to Claude Opus 5.5 for reflection, with settings, contexts, caching, and usage tracking.
- techniqueLLM-as-judge metricsWhen there is no single right answer: using a model with a rubric as the metric, writing the judge as a DSPy program, and keeping it honest.
Why you need this
GEPA: reflective prompt evolution explains the algorithm. This page is about operating it: which arguments to set, how many model calls a run will spend, how to read what comes back, and how to point the same optimizer at text that is not a DSPy program at all. Every argument name below was checked against the installed DSPy 3.4.0 and gepa 0.1.4 source.
The idea
A GEPA run needs four things from you: a metric that explains its scores, a reflection model, a budget, and two splits.
The metric with feedback
dspy.GEPA calls your metric with five arguments, (gold, pred, trace, pred_name, pred_trace), and the constructor raises a TypeError if your function cannot accept five. A three-argument metric written for BootstrapFewShot therefore fails; to share one metric across optimizers, define it as metric(gold, pred, trace=None, pred_name=None, pred_trace=None). gold is the labeled example, pred the program's output, trace the record of every predictor's inputs and outputs during the rollout (one run of the program on one example), pred_name the predictor GEPA is currently improving, and pred_trace that predictor's part of the trace. Return dspy.Prediction(score=..., feedback=...) with a score in [0, 1]. If you return a bare number, GEPA writes the feedback for you: This trajectory got a score of 0.0. That sentence tells the reflection model nothing, which is why feedback text matters. An LLM-as-judge metric that returns a justification is a ready-made feedback source.
The reflection model
reflection_lm is required (unless you supply a custom instruction_proposer). It reads failed rollouts and writes new instructions, so it should be strong. In this cluster it is Claude Opus 5.5, dspy.LM("anthropic/claude-opus-5-5", temperature=1.0, max_tokens=32000), with the key in ANTHROPIC_API_KEY. Those two settings follow the dspy.GEPA source, which recommends temperature 1.0 and max_tokens=32000 for the reflection model so it has room to write a long instruction. The program itself runs on the local task model served by Ollama, set with dspy.configure.
The budget rule
Exactly one of auto, max_full_evals, or max_metric_calls must be set, or the constructor fails an assertion. All three become a number of metric calls, and one metric call is one rollout:
| You set | GEPA uses |
|---|---|
max_metric_calls=B |
|
max_full_evals=F |
|
auto="light", "medium", "heavy" |
auto_budget, from candidates, the number of predictors, and the valset size |
auto_budget reuses MIPROv2's trial formula and adds up a baseline evaluation, bootstrap allowances, 35-example minibatches per trial, and periodic full valset evaluations. Calling it with a 50-example valset gives:
| Program | light | medium | heavy |
|---|---|---|---|
| 1 predictor | 580 | 940 | 1385 |
| 2 predictors | 1030 | 1390 | 1645 |
The budget is checked before each iteration, so a run can end up to one iteration past it.
Splits
compile(student, trainset=..., valset=...). The trainset feeds the reflection minibatches (3 examples each, reflection_minibatch_size=3); the valset is where every accepted candidate is scored for the Pareto front. Leave out the valset and GEPA uses the trainset for both and logs a warning that the prompts will overfit to it. The source also suggests a valset no bigger than needed to represent your real inputs, so more of the budget goes to exploring. teacher is not supported.
The other knobs
| Argument | Default | What it controls |
|---|---|---|
candidate_selection_strategy |
"pareto" |
Parent sampled from candidates that win on some valset example (the algorithm page explains the pruning), or always the best average with "current_best" |
component_selector |
"round_robin" |
Which predictor is rewritten each iteration; "all" rewrites all at once |
use_merge, max_merge_invocations |
True, 5 |
Merging two candidates that improved different predictors |
skip_perfect_score |
True |
Skip reflection when the whole minibatch already scores perfect_score |
log_dir |
None |
Saves every candidate; rerunning with the same directory resumes |
track_stats, track_best_outputs |
False |
Attach detailed_results (and the best output per valset example) |
num_threads, seed |
None, 0 |
Parallel rollouts; reproducibility |
instruction_proposer |
None |
Replace the default reflection prompt with your own proposer |
Reading the results
With track_stats=True, the returned program carries detailed_results:
| Field | Meaning |
|---|---|
candidates |
Every accepted program, index 0 being your original |
val_aggregate_scores |
Average valset score per candidate |
val_subscores |
Per-candidate dict of valset example index to score, which is |
per_val_instance_best_candidates |
For each valset example, the candidates that score best on it: the Pareto front |
parents |
Which candidate each one was mutated or merged from |
discovery_eval_counts |
Metric calls spent before each candidate was found |
best_idx, best_candidate |
Highest average score, and that program |
total_metric_calls, num_full_val_evals |
What the run actually spent |
If you pass your unlabeled batch of real tasks as the valset with track_best_outputs=True, best_outputs_valset holds the best output found for each one. That is GEPA as an inference-time search: you keep the answers, not the prompt.
Worked example
A full run (needs a real model; output shown is an illustration)
import dspy
LABELS = ["billing", "bug", "account"]
def metric(gold, pred, trace=None, pred_name=None, pred_trace=None):
got = (pred.category or "").strip()
if got == gold.category:
return dspy.Prediction(score=1.0, feedback="Correct.")
if got.lower() not in LABELS:
why = f"'{got}' is not a label. Answer with exactly one of: {', '.join(LABELS)}."
else:
why = f"Wrong label: answered '{got}', expected '{gold.category}'."
return dspy.Prediction(score=0.0, feedback=why)
dspy.configure(lm=dspy.LM("ollama_chat/llama3.2", api_base="http://localhost:11434", api_key=""))
reflection = dspy.LM("anthropic/claude-opus-5-5", temperature=1.0, max_tokens=32000)
optimizer = dspy.GEPA(metric=metric, auto="light", reflection_lm=reflection,
track_stats=True, log_dir="runs/triage-1", num_threads=8)
optimized = optimizer.compile(dspy.Predict("ticket -> category"),
trainset=trainset, valset=valset) # e.g. 100 and 50 examples
r = optimized.detailed_results
print(len(r.candidates), r.val_aggregate_scores[0], r.val_aggregate_scores[r.best_idx])
optimized.save("triage.json")
An illustration of the last print, not a real run: 7 0.62 0.84, which you would report as "62% to 84% on a 50-example valset", then confirm on the testset.
The same plumbing on DummyLM (run)
This file wires the same metric into dspy.GEPA with DummyLM standing in for both models. The fake task model answers each ticket from a fixed table, whatever the instruction says, so no proposal can help; watch GEPA notice that. Run with DSPy 3.4.0 and gepa 0.1.4 on Python 3.14 (3.12 and newer behave the same):
"""dspy.GEPA wired end to end on DummyLM: real plumbing, staged answers."""
import dspy
from dspy.utils import DummyLM
LABELS = ["billing", "bug", "account"]
def metric(gold, pred, trace=None, pred_name=None, pred_trace=None):
got = (pred.category or "").strip()
if got == gold.category:
return dspy.Prediction(score=1.0, feedback="Correct.")
if got.lower() not in LABELS:
why = f"'{got}' is not a label. Answer with exactly one of: {', '.join(LABELS)}."
else:
why = f"Wrong label: answered '{got}', expected '{gold.category}'."
return dspy.Prediction(score=0.0, feedback=why)
def ex(ticket, category):
return dspy.Example(ticket=ticket, category=category).with_inputs("ticket")
trainset = [
ex("I was charged twice this month.", "billing"),
ex("The app freezes on the login screen.", "bug"),
ex("Please close my account.", "account"),
ex("My card was declined but I was still billed.", "billing"),
]
valset = [
ex("Refund my last payment, please.", "billing"),
ex("Clicking Save does nothing.", "bug"),
ex("How do I change my username?", "account"),
ex("The PDF export is blank.", "bug"),
]
# The fake task model gives each ticket a fixed answer, whatever the instruction says.
task_lm = DummyLM({
"charged twice": {"category": "billing"},
"freezes": {"category": "Bug report"},
"close my account": {"category": "account"},
"declined": {"category": "billing"},
"Refund": {"category": "billing"},
"Save does nothing": {"category": "Bug report"},
"username": {"category": "account"},
"PDF export": {"category": "bug"},
})
FENCE = "`" * 3 # the reflection model puts its new instruction between fences
proposal = (
FENCE + "\nRead the support ticket and reply with exactly one lowercase label: "
"billing (charges, refunds, payments), bug (the product misbehaves), "
"or account (login details, closing or changing an account). No other words.\n" + FENCE
)
reflection_lm = DummyLM([{"new_instruction": proposal}] * 10)
dspy.configure(lm=task_lm)
program = dspy.Predict("ticket -> category")
optimizer = dspy.GEPA(
metric=metric,
max_metric_calls=20,
reflection_lm=reflection_lm,
num_threads=1,
track_stats=True,
seed=0,
)
optimized = optimizer.compile(program, trainset=trainset, valset=valset)
r = optimized.detailed_results
print("candidates:", len(r.candidates))
print("val_aggregate_scores:", r.val_aggregate_scores)
print("val_subscores[0]:", r.val_subscores[0])
print("total_metric_calls:", r.total_metric_calls)
print("reflection calls:", len(reflection_lm.history))
print("best instruction:", r.best_candidate.signature.instructions)
Run with python gepa_dummy.py. Output with timestamps and progress bars trimmed:
Running GEPA for approx 20 metric calls of the program. This amounts to 2.50 full evals on the train+val set.
Using 4 examples for tracking Pareto scores.
Iteration 0: Base program full valset score: 0.75 over 4 / 4 examples
Iteration 1: Selected program 0 score: 0.75
Iteration 1: Proposed new text for self: Read the support ticket and reply with exactly one lowercase label: billing (charges, refunds, payments), bug (the product misbehaves), or account (login details, closing or changing an account). No other words.
Iteration 1: New subsample score 2.0 is not better than old score 2.0, skipping
Iteration 2: Selected program 0 score: 0.75
Iteration 2: All subsample scores perfect for parent 0. Skipping.
Iteration 2: Reflective mutation did not propose a new candidate
Iteration 3: Selected program 0 score: 0.75
...
Iteration 4: New subsample score 2.0 is not better than old score 2.0, skipping
candidates: 1
val_aggregate_scores: [0.75]
val_subscores[0]: {0: 1.0, 1: 0.0, 2: 1.0, 3: 1.0}
total_metric_calls: 25
reflection calls: 3
best instruction: Given the fields `ticket`, produce the fields `category`.
What the run shows, line by line:
| Metric calls so far | Event |
|---|---|
| 4 | Baseline: 3 of 4 valset tickets right, 75% on a 4-example valset |
| 10 | Minibatch of 3 on the parent (2 right), reflection, same minibatch on the child (2 right): rejected, a tie is not an improvement |
| 13 | Minibatch all correct: skip_perfect_score skips reflection |
| 19, 25 | Two more proposals, both tied and rejected |
The run spent 25 calls against a budget of 20, because the check at 19 allowed one more iteration. The single-predictor program is named self, which is the name GEPA uses for the component. The val_subscores line is for the original program: it fails only example 1, the "Bug report" answer.
The standalone gepa library
The gepa package, which DSPy installs, optimizes any text you can score. Its entry point is optimize_anything; import it from the submodule, because the package attribute gepa.optimize_anything is the module, not the function. Your evaluator returns a score, or a score and a dict of diagnostics the reflection model reads (gepa calls this side information). With no dataset it runs single-task search, where the candidate is the answer; pass dataset (and valset) to optimize something that must work across many inputs. The reflection model can be a model-name string, which gepa passes to LiteLLM, or any function that takes a prompt and returns text. Here that function is scripted, so this complete file ran with no network:
"""gepa's optimize_anything on a regular expression, with a scripted reflection model."""
import re
from gepa.optimize_anything import EngineConfig, GEPAConfig, ReflectionConfig, optimize_anything
CASES = [ # (text, should it match?)
("TCK-1234", True),
("tck-0042", True),
("TCK-12", False),
("TCK-12345", False),
("TICKET-1234", False),
]
def evaluate(candidate: str):
failures = []
for text, want in CASES:
got = re.fullmatch(candidate, text) is not None
if got != want:
failures.append(f"{text!r}: matched={got}, expected {want}")
score = 1 - len(failures) / len(CASES)
return score, {"failures": failures or ["none"]}
def scripted_reflection(prompt):
# Stands in for the reflection model (Claude Opus 5.5). No network call.
scripted_reflection.calls += 1
fence = "`" * 3 # the new candidate goes between fences
return fence + "\n(?i)TCK-\\d{4}\n" + fence
scripted_reflection.calls = 0
result = optimize_anything(
seed_candidate=r"TCK-\d+",
evaluator=evaluate,
objective="A regex that fully matches ticket ids: TCK, a hyphen, exactly four digits; any letter case.",
config=GEPAConfig(
engine=EngineConfig(max_metric_calls=4),
reflection=ReflectionConfig(reflection_lm=scripted_reflection),
),
)
print("candidates:", result.candidates)
print("scores:", result.val_aggregate_scores)
print("best:", result.best_candidate)
print("metric calls:", result.total_metric_calls, "reflection calls:", scripted_reflection.calls)
Run with python anything_regex.py. Output, trimmed:
Iteration 0: Base program full valset score: 0.4 over 1 / 1 examples
Iteration 1: Selected program 0 score: 0.4
Iteration 1: Proposed new text for current_candidate: (?i)TCK-\d{4}
Iteration 1: Accepted candidate (subsample score 0.4 -> 1.0); running full eval.
Iteration 1: Found a better program on the valset with score 1.0.
...
candidates: [{'current_candidate': 'TCK-\\d+'}, {'current_candidate': '(?i)TCK-\\d{4}'}]
scores: [0.4, 1.0]
best: (?i)TCK-\d{4}
metric calls: 4 reflection calls: 1
The seed regex got 2 of 5 cases right (0.4): it missed the lowercase id and accepted 2 and 5 digits. The list of failures was the feedback. With a real reflection model you would pass reflection_lm="anthropic/claude-opus-5-5" in ReflectionConfig; the scripted reply here only shows the loop accepting an improvement.
In an optimization pipeline
This is the end of stage 3. Before a run: a baseline from Evaluating programs, a feedback metric, and separate trainset, valset, and testset. Start with auto="light" and a log_dir, read detailed_results, check the winner on the testset, and save it for stage 5. If the best candidate's valset gain does not survive the testset, see Overfitting and re-optimizing. For call-heavy deployments, Prompts and weights together covers following GEPA with fine-tuning.
Common mistakes
- Score-only feedback. A metric returning a bare float gives the reflection model "This trajectory got a score of 0.0." Symptom: proposals that reword the instruction without fixing the failure.
- A valset that is too small. Every accepted candidate is judged on it, so with 10 examples one lucky answer moves the score by 10 points. Symptom: a winner that loses on the testset.
- No valset at all. GEPA then tunes on the same examples it is judged on. Symptom: a near-perfect valset score and an ordinary testset score.
- A reflection model that is too weak. Symptom: proposals that ignore the feedback or repeat the old instruction.
- Setting two budget arguments. The constructor stops with "Exactly one of max_metric_calls, max_full_evals, auto must be set."
- Forgetting
log_diron a long run. An interrupted run starts over and spends the budget again.
Cost
Let be the budget in metric calls, the number of predictors (task-model calls per rollout), the valset size, the reflection minibatch size (3), and the input and output tokens per task-model call and per reflection call. Task-model tokens are about . A rejected iteration spends rollouts around one reflection call (6 at ); an accepted one adds a full valset evaluation, (56 at ). Reflection calls are therefore at most about , reached only if every proposal is rejected. Dollars are then
where and are your providers' prices per token for the task model and the reflection model. With a local task model the first term is zero dollars and becomes time instead: divided by num_threads, with seconds per call. Example: one predictor, 50-example valset, auto="light", so and . If, as an illustration, each reflection prompt is 3,000 tokens in and 400 out, that is at most 264,000 input and 35,200 output tokens of the reflection model, which you multiply by its current prices.
Going further
- The GEPA paper by Agrawal and colleagues (2025), for the experiments behind the defaults.
- The DSPy GEPA overview and "GEPA Advanced" documentation pages, for custom instruction proposers and the
MultiModalInstructionProposerfor image inputs. - The source of
dspy/teleprompt/gepa/gepa.py, especiallyauto_budgetand thecompilemethod. - The
gepapackage'soptimize_anythingdocstring, which describes single-task, multi-task, and generalization modes and a seedless mode where the reflection model writes the first candidate. - Saving and serving programs, for what to do with the winner.
Leads to
- techniquePrompts and weights togetherCombining prompt optimization and fine-tuning: alternating them with BetterTogether, ordering choices, and deciding which lever to pull first.
- techniqueSaving and serving programsKeeping an optimized program: saving state versus the whole program, loading it in production, pinning models and versions, and serving it behind an API.
Back to DSPy and GEPA: programming and optimizing language-model systems