technique
Writing metrics
A metric is the function an optimizer climbs: exact match, partial credit, multi-part scores, and the feedback text GEPA reads to learn why an output was wrong.
Before this
This page assumes you are comfortable with:
Why you need this
An optimizer does not know what you want. It knows one number per example, the one your metric returns, and it changes the program until that number goes up. A metric that rewards the wrong thing produces a program that does the wrong thing very reliably. For GEPA the metric does a second job: its feedback text is what the reflection model reads to decide how to rewrite an instruction, so a better-written explanation means better rewrites.
The idea
A metric is a plain Python function. DSPy calls it once per example with the gold example (your dspy.Example, with its labels) and the program's prediction, and it returns a score. In this cluster a score is a bool or a number between 0 and 1.
def metric(gold, pred, trace=None):
...
The three arguments, as DSPy 3.4.0 passes them:
| Argument | What it holds | When it is set |
|---|---|---|
gold |
The example: gold.answer, gold.kind, any label |
Always |
pred |
The program's dspy.Prediction: pred.answer |
Always |
trace |
The record of every predictor's inputs and outputs during this run | None when dspy.Evaluate scores; a list when an optimizer such as dspy.BootstrapFewShot is deciding whether to keep the run as demos |
The DSPy metrics documentation calls trace the lever that lets one metric behave two ways: a fine-grained score during evaluation, and a strict pass or fail during optimization, so only clean runs become worked examples.
Three return types
| Return | Meaning | How dspy.Evaluate totals it |
|---|---|---|
bool |
Right or wrong | Percentage of True |
float in [0, 1] |
Partial credit | Mean, shown as a percentage |
dspy.Prediction(score=..., feedback=...) |
A score plus an explanation in words | Mean of the scores; only GEPA reads the feedback |
Partial credit gives a number between 0 and 1 for answers that are partly right: the right value inside a long sentence, two of three required facts, a correct label with the wrong confidence. Combining checks is a weighted sum: if correctness matters three times as much as brevity, score = 0.75 * correct + 0.25 * short, where each check is 0 or 1 and the weights add to 1 so the result stays in [0, 1]. With correct = 1 and short = 0, the score is .
The GEPA feedback metric
dspy.GEPA needs more. In the installed source (dspy/teleprompt/gepa/gepa.py, protocol GEPAFeedbackMetric) the metric is called with five arguments:
def metric(gold, pred, trace=None, pred_name=None, pred_trace=None):
...
return dspy.Prediction(score=score, feedback=feedback)
| Argument | Meaning in the source |
|---|---|
gold, pred, trace |
As above |
pred_name |
The name of the predictor GEPA is improving right now, from named_predictors() |
pred_trace |
The part of the trace for that one predictor |
Facts checked in the 3.4.0 source:
- The
dspy.GEPAconstructor tests that the metric can be called with five positional arguments and raises aTypeErrorif not, so a three-argument(gold, pred, trace)metric is refused. A metric meant for GEPA and for everything else should be defined as(gold, pred, trace=None, pred_name=None, pred_trace=None): the defaults let the same function work withdspy.Evaluate, which calls it with onlygoldandpred, and withdspy.BootstrapFewShot, which passesgold,pred, andtrace. - When GEPA wants feedback, it calls the metric with
pred_nameandpred_tracefilled in, and readsscoreandfeedbackfrom the returneddspy.Prediction. - If the metric returns a plain number, or feedback of
None, GEPA writes the feedback itself:This trajectory got a score of {score}.That sentence tells the reflection model nothing about what went wrong. - GEPA uses the program-level score for selection. If the score returned with a
pred_namediffers, it logs a warning and ignores it; per-predictor scores are not supported yet, per-predictor feedback text is. - The protocol also has an optional sixth argument,
program_trace, which is filled in only when adspy.Flexsubmodule is being optimized.
Useful feedback says what was wrong, why, and what a right answer looks like, in terms the reflection model can turn into an instruction. "Wrong" is useless. "Expected 'Canberra', got 'Sydney'; the question asks for a capital, and the largest city is not always the capital" is something a rewritten instruction can act on.
Worked example
One question, three predictions, one metric written three ways. Run on DSPy 3.4.0 (on Python 3.14; 3.12 and newer behave the same); no model is called, because the predictions are written by hand.
# three_metrics.py
import dspy
def normalize(text):
return " ".join(text.lower().strip(" .").split())
# 1. Bool: right or wrong.
def exact_match(gold, pred, trace=None):
return normalize(pred.answer) == normalize(gold.answer)
# 2. Float: partial credit for the right answer buried in a long reply.
def graded(gold, pred, trace=None):
answer, target = normalize(pred.answer), normalize(gold.answer)
if answer == target:
score = 1.0
elif target in answer:
score = 0.5
else:
score = 0.0
if trace is not None: # inside an optimizer: only accept perfect demos
return score == 1.0
return score
# 3. Score plus feedback, in the five-argument form dspy.GEPA requires.
def graded_with_feedback(gold, pred, trace=None, pred_name=None, pred_trace=None):
score = graded(gold, pred)
answer, target = normalize(pred.answer), normalize(gold.answer)
if score == 1.0:
feedback = "Correct and in the expected short form."
elif score == 0.5:
extra = len(answer.split()) - len(target.split())
feedback = (f"The answer contains the right value '{gold.answer}' but adds {extra} "
"extra words. Reply with the value only, no sentence around it.")
else:
feedback = (f"Wrong: expected '{gold.answer}', got '{pred.answer}'. "
f"The question asks for a {gold.kind} fact; check it before answering.")
return dspy.Prediction(score=score, feedback=feedback)
if __name__ == "__main__":
gold = dspy.Example(question="What is the capital of Australia?", answer="Canberra",
kind="geography").with_inputs("question")
preds = [
dspy.Prediction(answer="Canberra."),
dspy.Prediction(answer="The capital of Australia is Canberra."),
dspy.Prediction(answer="Sydney"),
]
for p in preds:
print(repr(p.answer))
print(" bool :", exact_match(gold, p))
print(" float:", graded(gold, p), "| with a trace:", graded(gold, p, trace=[]))
fb = graded_with_feedback(gold, p)
print(" score:", fb.score)
print(" feedback:", fb.feedback)
Run with tmp/dspy-venv/Scripts/python.exe three_metrics.py:
'Canberra.'
bool : True
float: 1.0 | with a trace: True
score: 1.0
feedback: Correct and in the expected short form.
'The capital of Australia is Canberra.'
bool : False
float: 0.5 | with a trace: False
score: 0.5
feedback: The answer contains the right value 'Canberra' but adds 5 extra words. Reply with the value only, no sentence around it.
'Sydney'
bool : False
float: 0.0 | with a trace: False
score: 0.0
feedback: Wrong: expected 'Canberra', got 'Sydney'. The question asks for a geography fact; check it before answering.
The same three predictions, side by side:
| Prediction | bool | float | float with a trace | Feedback says |
|---|---|---|---|---|
Canberra. |
True | 1.0 | True | Correct |
The capital of Australia is Canberra. |
False | 0.5 | False | Right value, 5 extra words, drop the sentence |
Sydney |
False | 0.0 | False | Wrong value, and which kind of fact to check |
Normalizing (lower case, trimmed period, single spaces) is why Canberra. counts as exact. The "5 extra words" is computed: the normalized answer has 6 words and the target 1. With a trace, the partial-credit answer becomes False, so a bootstrapping optimizer would not copy a wordy answer into its demos.
A second script checks the five-argument rule and that the feedback metric also works with dspy.Evaluate. DummyLM is DSPy's fake LM that returns scripted answers; the reflection model here is a placeholder that is never called. In real use the task model is the local task model (a model served by Ollama) and the reflection model is Claude Opus 5.5.
# gepa_accepts.py
import dspy
from dspy.utils import DummyLM
from three_metrics import exact_match, graded_with_feedback
reflection = DummyLM([])
try:
dspy.GEPA(metric=exact_match, auto="light", reflection_lm=reflection)
except TypeError as e:
print("TypeError:", str(e).split(". See")[0])
gepa = dspy.GEPA(metric=graded_with_feedback, auto="light", reflection_lm=reflection)
print("accepted:", type(gepa).__name__)
devset = [
dspy.Example(question="What is the capital of Australia?", answer="Canberra",
kind="geography").with_inputs("question"),
dspy.Example(question="What is 12 times 12?", answer="144",
kind="math").with_inputs("question"),
]
dspy.configure(lm=DummyLM([{"answer": "Canberra"}, {"answer": "It is 144."}]))
evaluate = dspy.Evaluate(devset=devset, metric=graded_with_feedback, num_threads=1)
result = evaluate(dspy.Predict("question -> answer"))
print("score:", result.score)
Output:
2026/10/03 09:15:07 INFO dspy.evaluate.evaluate: Average Metric: 1.5 / 2 (75.0%)
TypeError: GEPA metric must accept five arguments: (gold, pred, trace, pred_name, pred_trace)
accepted: GEPA
score: 75.0
The three-argument metric is refused when GEPA is built, before any model call. The feedback metric scored 1.0 and 0.5, so dspy.Evaluate reports on this 2-example devset. (The log line prints first because logging and printing go to different streams.)
In an optimization pipeline
This is stage 2, measuring the program, and it steers all of stage 3. dspy.BootstrapFewShot uses the metric with a trace to decide which runs become demos. dspy.MIPROv2 uses the plain score to rank instruction and demo combinations. dspy.GEPA uses the score to keep candidates on its per-example Pareto front and the feedback text to write new instructions. A metric that is wrong in a subtle way is wrong in every one of those places at once.
Common mistakes
- Score-only feedback for GEPA. A plain number becomes "This trajectory got a score of 0.0." Symptom: GEPA's proposed instructions are vague restatements of the task.
- Metrics that are easy to game. "Score 1 if the gold answer appears anywhere in the reply" rewards a reply that lists every capital in the world. Symptom: scores climb while replies get longer.
- Returning a number outside [0, 1]. A 0 to 10 rubric score breaks
failure_score=0.0andperfect_score=1.0assumptions in GEPA. Divide by the maximum. - Crashing on odd output.
pred.answer.lower()fails if the answer isNoneafter a parse failure. Guard with(pred.answer or ""). - Feedback that leaks the answer for every example. Telling the reflection model each gold answer invites instructions that memorize valset quirks. Explain the kind of mistake; quoting the expected value for one example is fine, building a lookup table is not.
Cost
A rule-based metric like the ones above costs microseconds per call, and an optimizer calls it once per rollout, so a GEPA budget of metric calls (its max_metric_calls) means metric calls and about program runs. The cost appears when the metric itself calls a model, as an LLM judge does: then each metric call adds judge calls, and a run costs about dollars, with predictors per program run and the average cost of one call of each kind. Builder time is the real expense: a good feedback metric takes an hour or two to write and a few rounds of reading GEPA's proposals to tune.
Going further
- Building a dataset: the examples whose labels the metric reads.
- LLM-as-judge metrics: when no rule can score the output, and the judge's justification doubles as GEPA feedback.
- Evaluating programs: running the metric over a whole set.
- GEPA: reflective prompt evolution: what the reflection model does with your feedback text.
- The
GEPAFeedbackMetricdocstring in the installeddspy.GEPAsource.
Leads to
- techniqueEvaluating programsRunning a program over a dataset with dspy.Evaluate, reading the results, and deciding whether a change is a real improvement or noise.
- techniqueGEPA: reflective prompt evolutionHow GEPA improves prompts by reading what went wrong: run a minibatch, reflect on traces and feedback in plain language, propose a new instruction, and keep a Pareto front of candidates.
- techniqueLLM-as-judge metricsWhen there is no single right answer: using a model with a rubric as the metric, writing the judge as a DSPy program, and keeping it honest.
Back to DSPy and GEPA: programming and optimizing language-model systems