technique
Evaluating programs
Running a program over a dataset with dspy.Evaluate, reading the results, and deciding whether a change is a real improvement or noise.
Before this
This page assumes you are comfortable with:
- techniqueWriting metricsA metric is the function an optimizer climbs: exact match, partial credit, multi-part scores, and the feedback text GEPA reads to learn why an output was wrong.
- prerequisiteNoisy scores and sample sizeWhy a score measured on a small set wobbles, how to estimate the wobble, and how many examples you need before a difference is real.
Why you need this
Every optimizer in stage 3 and stage 4 makes one claim: "the new program is better than the old one." Evaluation is how you check that claim. A program is a DSPy module tree; a metric is a function that scores one output. Evaluating a program means running it on every example in a set, scoring each output with the metric, and averaging. If you skip this step, or do it carelessly, you cannot tell an improvement from luck, and every number later in the pipeline inherits the doubt.
The idea
In plain words: run the program on a fixed list of examples, score each answer, and report the average with the name and size of the list. Then look at the individual rows, because the average hides which examples failed and why.
The precise version uses this cluster's notation. Write for the score of program on example , a number in returned by the metric. On a set of examples the program's score is the average
With small numbers: a program that gets 8 of 12 examples right has , which DSPy reports as 66.67%.
DSPy does this with dspy.Evaluate. In DSPy 3.4 (the version the sample on this page ran against, on Python 3.14; 3.12 and newer behave the same) its constructor takes keyword arguments only:
| Argument | What it does |
|---|---|
devset |
The list of dspy.Example objects to run. DSPy's own name for the evaluation set, whichever split you pass. |
metric |
Your metric function, called as metric(example, prediction). |
num_threads |
How many examples run at once. More threads finish sooner but send calls in parallel. |
display_progress, display_table |
A progress bar, and a table of rows (the table needs the pandas package; without it DSPy logs a warning and skips it). |
max_errors |
Stop after this many crashed examples. |
failure_score |
The score given to an example whose run raised an error; defaults to 0.0. |
save_as_csv, save_as_json |
Write every row to a file for later analysis. |
Calling the evaluator on a program returns an EvaluationResult. Its score is a percentage rounded to two decimals, and its results is a list of (example, prediction, score) tuples, one per example, which is where per-example inspection starts.
Worked example
The task is two-label sentiment on twelve short product reviews. To make the score known in advance, the language model is DSPy's fake model, DummyLM, given a scripted answer for each review. Two scripts stand in for two versions of the program (say, two different instructions): version A gets reviews 3, 4, 9, and 10 wrong; version B gets 6 and 10 wrong. No real model is called.
import dspy
from dspy.utils import DummyLM
reviews = [
("Arrived fast and works perfectly.", "positive"),
("Broke after two days.", "negative"),
("Not bad at all, I would buy it again.", "positive"),
("The color is nice but it does not charge.", "negative"),
("Five stars, my kids love it.", "positive"),
("I expected more for the price.", "negative"),
("Could not be happier.", "positive"),
("Returned it the same week.", "negative"),
("Not what I hoped, the strap snapped.", "negative"),
("Hard to set up, but great once running.", "positive"),
("Never again.", "negative"),
("Exactly as described.", "positive"),
]
devset = [dspy.Example(text=t, label=y).with_inputs("text") for t, y in reviews]
# Scripted answers, keyed by review text, so each program's mistakes are known in advance.
# Program A misses rows 3, 4, 9, 10. Program B misses rows 6 and 10.
a_wrong = {2, 3, 8, 9}
b_wrong = {5, 9}
flip = {"positive": "negative", "negative": "positive"}
def script(wrong):
return {t: {"label": flip[y] if i in wrong else y} for i, (t, y) in enumerate(reviews)}
lm_a = DummyLM(script(a_wrong))
lm_b = DummyLM(script(b_wrong))
program = dspy.Predict("text -> label")
def exact_label(example, pred, trace=None):
return example.label == pred.label
evaluate = dspy.Evaluate(devset=devset, metric=exact_label, num_threads=1)
with dspy.context(lm=lm_a):
result_a = evaluate(program)
with dspy.context(lm=lm_b):
result_b = evaluate(program)
print("A:", result_a)
print("A score:", result_a.score)
print("B score:", result_b.score)
print()
print("row gold A B")
for i, ((ex, pa, sa), (_, pb, sb)) in enumerate(zip(result_a.results, result_b.results), start=1):
print(f"{i:>3} {ex.label:<9} {int(sa)} {pa.label:<8} {int(sb)} {pb.label}")
only_a = sum(1 for (_, _, sa), (_, _, sb) in zip(result_a.results, result_b.results) if sa and not sb)
only_b = sum(1 for (_, _, sa), (_, _, sb) in zip(result_a.results, result_b.results) if sb and not sa)
print()
print("only A right:", only_a, " only B right:", only_b)
Run it with python evaluate_two.py. It printed:
2026/10/03 09:11:26 INFO dspy.evaluate.evaluate: Average Metric: 8 / 12 (66.7%)
2026/10/03 09:11:26 INFO dspy.evaluate.evaluate: Average Metric: 10 / 12 (83.3%)
A: EvaluationResult(score=66.67, results=<list of 12 results>)
A score: 66.67
B score: 83.33
row gold A B
1 positive 1 positive 1 positive
2 negative 1 negative 1 negative
3 positive 0 negative 1 positive
4 negative 0 positive 1 negative
5 positive 1 positive 1 positive
6 negative 1 negative 0 positive
7 positive 1 positive 1 positive
8 negative 1 negative 1 negative
9 negative 0 positive 1 negative
10 positive 0 negative 0 negative
11 negative 1 negative 1 negative
12 positive 1 positive 1 positive
only A right: 1 only B right: 3
The baseline. Version A scores 66.67% on the 12-example devset. That is the number every later change is compared against. Record it with the split name and size, the model, and the date.
Per-example inspection. Reading A's failures by row shows a pattern (the script put it there, but real failures cluster the same way): row 3 ("Not bad at all") and row 9 ("Not what I hoped") both start with "Not", and rows 4 and 10 both turn around after a "but". Grouping failures by a shared cause like this is error analysis. It tells you what to change next far better than the average does.
Is B really better? B scores 83.33%, which is 16.67 points higher. Before believing it, recall the standard error from Noisy scores and sample size: for a score on examples it is . For A that is , so a rough 95% range is plus or minus 27 points. For B it is about 0.108, plus or minus 22 points. The ranges overlap almost completely.
The paired comparison does better, because both versions ran on the same twelve reviews. Rows where both are right or both are wrong carry no information about which is better. Only the disagreements matter: B alone is right on 3 rows (3, 4, 9), A alone on 1 row (6). If the two versions were equally good, each disagreement would be a coin flip. The chance of a split at least as lopsided as 3 to 1, in either direction, is
A 62.5% chance of seeing this by luck is not evidence. Four disagreements are too few. The honest conclusion is "B looks promising; collect more examples", not "B is 17 points better".
In an optimization pipeline
Evaluation is the last step of stage 2 and the referee for everything after it.
| Moment | Which split | Why |
|---|---|---|
| Before any optimizer runs | valset |
The baseline. Without it, "GEPA improved the program" has nothing to compare against. |
| Inside an optimizer | trainset and valset |
Optimizers call the same metric thousands of times. GEPA, for example, scores every accepted candidate on the whole valset. |
| After optimizing | valset, then testset once |
The valset was used to choose, so its score is flattering. The testset gives the number you report. |
| After a model or data change | the same splits as before | Same examples, same metric, so the comparison is paired. |
Two practical rules. Keep the evaluation set fixed between versions, so every comparison is paired. And save the rows (save_as_json), not just the average, so you can rerun error analysis later without paying for new calls.
Common mistakes
- Reporting a bare percentage. "83%" with no split or size hides that it is 10 of 12. Symptom: a later run on 200 examples "drops" to 74% and nobody can say whether anything changed.
- Rerunning to measure noise while the cache is on. DSPy caches language-model responses by default, so a second run with identical inputs returns the identical answers and the identical score. Symptom: two runs agree to the decimal, which says nothing about sampling variation. To get fresh samples, use a nonzero temperature and a different
rollout_idper run (DSPy 3.4 ignoresrollout_idat temperature 0). - Comparing on different examples. A on last week's devset, B on this week's. Symptom: a 10-point swing that disappears when both run on the same list.
- Letting crashes count as wrong answers silently. An example that raises an error gets
failure_score(0.0 by default). Symptom: a score drop that is really a parsing bug; the rows show empty predictions. - Evaluating on the set you tuned on. Symptom: valset score climbs every week while users see no change. Keep a testset that no optimizer and no human tweak has touched.
- Threads that hit a rate limit. A high
num_threadsagainst a hosted model can trip the provider's limits and turn examples into errors. Symptom: the score depends on how busy the service was.
Cost
One evaluation is one rollout per example, so with examples and a program of predictors it makes model calls, plus the metric's own calls if the metric uses a model as a judge (see LLM-as-judge metrics). In dollars, if each call uses input tokens and output tokens at prices and per token, one evaluation costs about . With the local task model the dollar cost is near zero and the price is wall-clock time: roughly times the time per call, divided by num_threads if the server can handle that many requests at once. The precision you buy grows only with : halving the standard error takes four times as many examples. Evaluation starts to hurt when the metric itself is a strong model, because then every example costs a strong-model call.
Going further
- Read LLM-as-judge metrics for metrics that need a model to grade open-ended answers.
- Try Few-shot bootstrapping on a program you have a baseline for, and compare the two with a paired count of disagreements.
- McNemar's test, the standard name for the paired disagreement test used above.
- Bootstrap confidence intervals: resample your rows with replacement to see how much the average moves.
- Overfitting and re-optimizing, for when the valset and testset disagree.
Leads to
- techniqueFew-shot bootstrappingThe first family of optimizers: LabeledFewShot, BootstrapFewShot, and random search over bootstrapped demos, which pick worked examples for every predictor automatically.
- techniqueOverfitting and re-optimizingKnowing when an optimized program has stopped being good: overfitting to the valset, the model changing underneath you, data drift, and when to run the optimizer again.
Back to DSPy and GEPA: programming and optimizing language-model systems