technique

Evaluating programs

Running a program over a dataset with dspy.Evaluate, reading the results, and deciding whether a change is a real improvement or noise.

Before this

This page assumes you are comfortable with:

Why you need this

Every optimizer in stage 3 and stage 4 makes one claim: "the new program is better than the old one." Evaluation is how you check that claim. A program is a DSPy module tree; a metric is a function that scores one output. Evaluating a program means running it on every example in a set, scoring each output with the metric, and averaging. If you skip this step, or do it carelessly, you cannot tell an improvement from luck, and every number later in the pipeline inherits the doubt.

The idea

In plain words: run the program on a fixed list of examples, score each answer, and report the average with the name and size of the list. Then look at the individual rows, because the average hides which examples failed and why.

The precise version uses this cluster's notation. Write si(p)s_i(p) for the score of program pp on example ii, a number in [0,1][0, 1] returned by the metric. On a set of nn examples the program's score is the average

sˉ(p)=1n∑i=1nsi(p).\bar{s}(p) = \frac{1}{n}\sum_{i=1}^{n} s_i(p).

With small numbers: a program that gets 8 of 12 examples right has sˉ(p)=8/12≈0.667\bar{s}(p) = 8/12 \approx 0.667, which DSPy reports as 66.67%.

DSPy does this with dspy.Evaluate. In DSPy 3.4 (the version the sample on this page ran against, on Python 3.14; 3.12 and newer behave the same) its constructor takes keyword arguments only:

Argument What it does
devset The list of dspy.Example objects to run. DSPy's own name for the evaluation set, whichever split you pass.
metric Your metric function, called as metric(example, prediction).
num_threads How many examples run at once. More threads finish sooner but send calls in parallel.
display_progress, display_table A progress bar, and a table of rows (the table needs the pandas package; without it DSPy logs a warning and skips it).
max_errors Stop after this many crashed examples.
failure_score The score given to an example whose run raised an error; defaults to 0.0.
save_as_csv, save_as_json Write every row to a file for later analysis.

Calling the evaluator on a program returns an EvaluationResult. Its score is a percentage rounded to two decimals, and its results is a list of (example, prediction, score) tuples, one per example, which is where per-example inspection starts.

Worked example

The task is two-label sentiment on twelve short product reviews. To make the score known in advance, the language model is DSPy's fake model, DummyLM, given a scripted answer for each review. Two scripts stand in for two versions of the program (say, two different instructions): version A gets reviews 3, 4, 9, and 10 wrong; version B gets 6 and 10 wrong. No real model is called.

import dspy
from dspy.utils import DummyLM

reviews = [
    ("Arrived fast and works perfectly.", "positive"),
    ("Broke after two days.", "negative"),
    ("Not bad at all, I would buy it again.", "positive"),
    ("The color is nice but it does not charge.", "negative"),
    ("Five stars, my kids love it.", "positive"),
    ("I expected more for the price.", "negative"),
    ("Could not be happier.", "positive"),
    ("Returned it the same week.", "negative"),
    ("Not what I hoped, the strap snapped.", "negative"),
    ("Hard to set up, but great once running.", "positive"),
    ("Never again.", "negative"),
    ("Exactly as described.", "positive"),
]
devset = [dspy.Example(text=t, label=y).with_inputs("text") for t, y in reviews]

# Scripted answers, keyed by review text, so each program's mistakes are known in advance.
# Program A misses rows 3, 4, 9, 10. Program B misses rows 6 and 10.
a_wrong = {2, 3, 8, 9}
b_wrong = {5, 9}
flip = {"positive": "negative", "negative": "positive"}

def script(wrong):
    return {t: {"label": flip[y] if i in wrong else y} for i, (t, y) in enumerate(reviews)}

lm_a = DummyLM(script(a_wrong))
lm_b = DummyLM(script(b_wrong))

program = dspy.Predict("text -> label")

def exact_label(example, pred, trace=None):
    return example.label == pred.label

evaluate = dspy.Evaluate(devset=devset, metric=exact_label, num_threads=1)

with dspy.context(lm=lm_a):
    result_a = evaluate(program)
with dspy.context(lm=lm_b):
    result_b = evaluate(program)

print("A:", result_a)
print("A score:", result_a.score)
print("B score:", result_b.score)
print()
print("row  gold      A         B")
for i, ((ex, pa, sa), (_, pb, sb)) in enumerate(zip(result_a.results, result_b.results), start=1):
    print(f"{i:>3}  {ex.label:<9} {int(sa)} {pa.label:<8} {int(sb)} {pb.label}")

only_a = sum(1 for (_, _, sa), (_, _, sb) in zip(result_a.results, result_b.results) if sa and not sb)
only_b = sum(1 for (_, _, sa), (_, _, sb) in zip(result_a.results, result_b.results) if sb and not sa)
print()
print("only A right:", only_a, " only B right:", only_b)

Run it with python evaluate_two.py. It printed:

2026/10/03 09:11:26 INFO dspy.evaluate.evaluate: Average Metric: 8 / 12 (66.7%)
2026/10/03 09:11:26 INFO dspy.evaluate.evaluate: Average Metric: 10 / 12 (83.3%)
A: EvaluationResult(score=66.67, results=<list of 12 results>)
A score: 66.67
B score: 83.33

row  gold      A         B
  1  positive  1 positive 1 positive
  2  negative  1 negative 1 negative
  3  positive  0 negative 1 positive
  4  negative  0 positive 1 negative
  5  positive  1 positive 1 positive
  6  negative  1 negative 0 positive
  7  positive  1 positive 1 positive
  8  negative  1 negative 1 negative
  9  negative  0 positive 1 negative
 10  positive  0 negative 0 negative
 11  negative  1 negative 1 negative
 12  positive  1 positive 1 positive

only A right: 1  only B right: 3

The baseline. Version A scores 66.67% on the 12-example devset. That is the number every later change is compared against. Record it with the split name and size, the model, and the date.

Per-example inspection. Reading A's failures by row shows a pattern (the script put it there, but real failures cluster the same way): row 3 ("Not bad at all") and row 9 ("Not what I hoped") both start with "Not", and rows 4 and 10 both turn around after a "but". Grouping failures by a shared cause like this is error analysis. It tells you what to change next far better than the average does.

Is B really better? B scores 83.33%, which is 16.67 points higher. Before believing it, recall the standard error from Noisy scores and sample size: for a score pp on nn examples it is p(1−p)/n\sqrt{p(1-p)/n}. For A that is 0.667×0.333/12≈0.136\sqrt{0.667 \times 0.333 / 12} \approx 0.136, so a rough 95% range is plus or minus 27 points. For B it is about 0.108, plus or minus 22 points. The ranges overlap almost completely.

The paired comparison does better, because both versions ran on the same twelve reviews. Rows where both are right or both are wrong carry no information about which is better. Only the disagreements matter: B alone is right on 3 rows (3, 4, 9), A alone on 1 row (6). If the two versions were equally good, each disagreement would be a coin flip. The chance of a split at least as lopsided as 3 to 1, in either direction, is

2×(43)+(44)24=2×4+116=0.625.2 \times \frac{\binom{4}{3} + \binom{4}{4}}{2^4} = 2 \times \frac{4 + 1}{16} = 0.625.

A 62.5% chance of seeing this by luck is not evidence. Four disagreements are too few. The honest conclusion is "B looks promising; collect more examples", not "B is 17 points better".

In an optimization pipeline

Evaluation is the last step of stage 2 and the referee for everything after it.

Moment Which split Why
Before any optimizer runs valset The baseline. Without it, "GEPA improved the program" has nothing to compare against.
Inside an optimizer trainset and valset Optimizers call the same metric thousands of times. GEPA, for example, scores every accepted candidate on the whole valset.
After optimizing valset, then testset once The valset was used to choose, so its score is flattering. The testset gives the number you report.
After a model or data change the same splits as before Same examples, same metric, so the comparison is paired.

Two practical rules. Keep the evaluation set fixed between versions, so every comparison is paired. And save the rows (save_as_json), not just the average, so you can rerun error analysis later without paying for new calls.

Common mistakes

  • Reporting a bare percentage. "83%" with no split or size hides that it is 10 of 12. Symptom: a later run on 200 examples "drops" to 74% and nobody can say whether anything changed.
  • Rerunning to measure noise while the cache is on. DSPy caches language-model responses by default, so a second run with identical inputs returns the identical answers and the identical score. Symptom: two runs agree to the decimal, which says nothing about sampling variation. To get fresh samples, use a nonzero temperature and a different rollout_id per run (DSPy 3.4 ignores rollout_id at temperature 0).
  • Comparing on different examples. A on last week's devset, B on this week's. Symptom: a 10-point swing that disappears when both run on the same list.
  • Letting crashes count as wrong answers silently. An example that raises an error gets failure_score (0.0 by default). Symptom: a score drop that is really a parsing bug; the rows show empty predictions.
  • Evaluating on the set you tuned on. Symptom: valset score climbs every week while users see no change. Keep a testset that no optimizer and no human tweak has touched.
  • Threads that hit a rate limit. A high num_threads against a hosted model can trip the provider's limits and turn examples into errors. Symptom: the score depends on how busy the service was.

Cost

One evaluation is one rollout per example, so with nn examples and a program of kk predictors it makes n×kn \times k model calls, plus the metric's own calls if the metric uses a model as a judge (see LLM-as-judge metrics). In dollars, if each call uses tint_{\text{in}} input tokens and toutt_{\text{out}} output tokens at prices cinc_{\text{in}} and coutc_{\text{out}} per token, one evaluation costs about nk(tincin+toutcout)n k (t_{\text{in}} c_{\text{in}} + t_{\text{out}} c_{\text{out}}). With the local task model the dollar cost is near zero and the price is wall-clock time: roughly nkn k times the time per call, divided by num_threads if the server can handle that many requests at once. The precision you buy grows only with n\sqrt{n}: halving the standard error takes four times as many examples. Evaluation starts to hurt when the metric itself is a strong model, because then every example costs a strong-model call.

Going further

  • Read LLM-as-judge metrics for metrics that need a model to grade open-ended answers.
  • Try Few-shot bootstrapping on a program you have a baseline for, and compare the two with a paired count of disagreements.
  • McNemar's test, the standard name for the paired disagreement test used above.
  • Bootstrap confidence intervals: resample your rows with replacement to see how much the average moves.
  • Overfitting and re-optimizing, for when the valset and testset disagree.

Leads to

Back to DSPy and GEPA: programming and optimizing language-model systems