technique

Writing metrics

A metric is the function an optimizer climbs: exact match, partial credit, multi-part scores, and the feedback text GEPA reads to learn why an output was wrong.

Before this

This page assumes you are comfortable with:

Why you need this

An optimizer does not know what you want. It knows one number per example, the one your metric returns, and it changes the program until that number goes up. A metric that rewards the wrong thing produces a program that does the wrong thing very reliably. For GEPA the metric does a second job: its feedback text is what the reflection model reads to decide how to rewrite an instruction, so a better-written explanation means better rewrites.

The idea

A metric is a plain Python function. DSPy calls it once per example with the gold example (your dspy.Example, with its labels) and the program's prediction, and it returns a score. In this cluster a score is a bool or a number between 0 and 1.

def metric(gold, pred, trace=None):
    ...

The three arguments, as DSPy 3.4.0 passes them:

Argument What it holds When it is set
gold The example: gold.answer, gold.kind, any label Always
pred The program's dspy.Prediction: pred.answer Always
trace The record of every predictor's inputs and outputs during this run None when dspy.Evaluate scores; a list when an optimizer such as dspy.BootstrapFewShot is deciding whether to keep the run as demos

The DSPy metrics documentation calls trace the lever that lets one metric behave two ways: a fine-grained score during evaluation, and a strict pass or fail during optimization, so only clean runs become worked examples.

Three return types

Return Meaning How dspy.Evaluate totals it
bool Right or wrong Percentage of True
float in [0, 1] Partial credit Mean, shown as a percentage
dspy.Prediction(score=..., feedback=...) A score plus an explanation in words Mean of the scores; only GEPA reads the feedback

Partial credit gives a number between 0 and 1 for answers that are partly right: the right value inside a long sentence, two of three required facts, a correct label with the wrong confidence. Combining checks is a weighted sum: if correctness matters three times as much as brevity, score = 0.75 * correct + 0.25 * short, where each check is 0 or 1 and the weights add to 1 so the result stays in [0, 1]. With correct = 1 and short = 0, the score is 0.75×1+0.25×0=0.750.75 \times 1 + 0.25 \times 0 = 0.75.

The GEPA feedback metric

dspy.GEPA needs more. In the installed source (dspy/teleprompt/gepa/gepa.py, protocol GEPAFeedbackMetric) the metric is called with five arguments:

def metric(gold, pred, trace=None, pred_name=None, pred_trace=None):
    ...
    return dspy.Prediction(score=score, feedback=feedback)
Argument Meaning in the source
gold, pred, trace As above
pred_name The name of the predictor GEPA is improving right now, from named_predictors()
pred_trace The part of the trace for that one predictor

Facts checked in the 3.4.0 source:

  • The dspy.GEPA constructor tests that the metric can be called with five positional arguments and raises a TypeError if not, so a three-argument (gold, pred, trace) metric is refused. A metric meant for GEPA and for everything else should be defined as (gold, pred, trace=None, pred_name=None, pred_trace=None): the defaults let the same function work with dspy.Evaluate, which calls it with only gold and pred, and with dspy.BootstrapFewShot, which passes gold, pred, and trace.
  • When GEPA wants feedback, it calls the metric with pred_name and pred_trace filled in, and reads score and feedback from the returned dspy.Prediction.
  • If the metric returns a plain number, or feedback of None, GEPA writes the feedback itself: This trajectory got a score of {score}. That sentence tells the reflection model nothing about what went wrong.
  • GEPA uses the program-level score for selection. If the score returned with a pred_name differs, it logs a warning and ignores it; per-predictor scores are not supported yet, per-predictor feedback text is.
  • The protocol also has an optional sixth argument, program_trace, which is filled in only when a dspy.Flex submodule is being optimized.

Useful feedback says what was wrong, why, and what a right answer looks like, in terms the reflection model can turn into an instruction. "Wrong" is useless. "Expected 'Canberra', got 'Sydney'; the question asks for a capital, and the largest city is not always the capital" is something a rewritten instruction can act on.

Worked example

One question, three predictions, one metric written three ways. Run on DSPy 3.4.0 (on Python 3.14; 3.12 and newer behave the same); no model is called, because the predictions are written by hand.

# three_metrics.py
import dspy


def normalize(text):
    return " ".join(text.lower().strip(" .").split())


# 1. Bool: right or wrong.
def exact_match(gold, pred, trace=None):
    return normalize(pred.answer) == normalize(gold.answer)


# 2. Float: partial credit for the right answer buried in a long reply.
def graded(gold, pred, trace=None):
    answer, target = normalize(pred.answer), normalize(gold.answer)
    if answer == target:
        score = 1.0
    elif target in answer:
        score = 0.5
    else:
        score = 0.0
    if trace is not None:  # inside an optimizer: only accept perfect demos
        return score == 1.0
    return score


# 3. Score plus feedback, in the five-argument form dspy.GEPA requires.
def graded_with_feedback(gold, pred, trace=None, pred_name=None, pred_trace=None):
    score = graded(gold, pred)
    answer, target = normalize(pred.answer), normalize(gold.answer)
    if score == 1.0:
        feedback = "Correct and in the expected short form."
    elif score == 0.5:
        extra = len(answer.split()) - len(target.split())
        feedback = (f"The answer contains the right value '{gold.answer}' but adds {extra} "
                    "extra words. Reply with the value only, no sentence around it.")
    else:
        feedback = (f"Wrong: expected '{gold.answer}', got '{pred.answer}'. "
                    f"The question asks for a {gold.kind} fact; check it before answering.")
    return dspy.Prediction(score=score, feedback=feedback)


if __name__ == "__main__":
    gold = dspy.Example(question="What is the capital of Australia?", answer="Canberra",
                        kind="geography").with_inputs("question")
    preds = [
        dspy.Prediction(answer="Canberra."),
        dspy.Prediction(answer="The capital of Australia is Canberra."),
        dspy.Prediction(answer="Sydney"),
    ]

    for p in preds:
        print(repr(p.answer))
        print("  bool :", exact_match(gold, p))
        print("  float:", graded(gold, p), "| with a trace:", graded(gold, p, trace=[]))
        fb = graded_with_feedback(gold, p)
        print("  score:", fb.score)
        print("  feedback:", fb.feedback)

Run with tmp/dspy-venv/Scripts/python.exe three_metrics.py:

'Canberra.'
  bool : True
  float: 1.0 | with a trace: True
  score: 1.0
  feedback: Correct and in the expected short form.
'The capital of Australia is Canberra.'
  bool : False
  float: 0.5 | with a trace: False
  score: 0.5
  feedback: The answer contains the right value 'Canberra' but adds 5 extra words. Reply with the value only, no sentence around it.
'Sydney'
  bool : False
  float: 0.0 | with a trace: False
  score: 0.0
  feedback: Wrong: expected 'Canberra', got 'Sydney'. The question asks for a geography fact; check it before answering.

The same three predictions, side by side:

Prediction bool float float with a trace Feedback says
Canberra. True 1.0 True Correct
The capital of Australia is Canberra. False 0.5 False Right value, 5 extra words, drop the sentence
Sydney False 0.0 False Wrong value, and which kind of fact to check

Normalizing (lower case, trimmed period, single spaces) is why Canberra. counts as exact. The "5 extra words" is computed: the normalized answer has 6 words and the target 1. With a trace, the partial-credit answer becomes False, so a bootstrapping optimizer would not copy a wordy answer into its demos.

A second script checks the five-argument rule and that the feedback metric also works with dspy.Evaluate. DummyLM is DSPy's fake LM that returns scripted answers; the reflection model here is a placeholder that is never called. In real use the task model is the local task model (a model served by Ollama) and the reflection model is Claude Opus 5.5.

# gepa_accepts.py
import dspy
from dspy.utils import DummyLM

from three_metrics import exact_match, graded_with_feedback

reflection = DummyLM([])

try:
    dspy.GEPA(metric=exact_match, auto="light", reflection_lm=reflection)
except TypeError as e:
    print("TypeError:", str(e).split(". See")[0])

gepa = dspy.GEPA(metric=graded_with_feedback, auto="light", reflection_lm=reflection)
print("accepted:", type(gepa).__name__)

devset = [
    dspy.Example(question="What is the capital of Australia?", answer="Canberra",
                 kind="geography").with_inputs("question"),
    dspy.Example(question="What is 12 times 12?", answer="144",
                 kind="math").with_inputs("question"),
]
dspy.configure(lm=DummyLM([{"answer": "Canberra"}, {"answer": "It is 144."}]))
evaluate = dspy.Evaluate(devset=devset, metric=graded_with_feedback, num_threads=1)
result = evaluate(dspy.Predict("question -> answer"))
print("score:", result.score)

Output:

2026/10/03 09:15:07 INFO dspy.evaluate.evaluate: Average Metric: 1.5 / 2 (75.0%)
TypeError: GEPA metric must accept five arguments: (gold, pred, trace, pred_name, pred_trace)
accepted: GEPA
score: 75.0

The three-argument metric is refused when GEPA is built, before any model call. The feedback metric scored 1.0 and 0.5, so dspy.Evaluate reports (1.0+0.5)/2=75%(1.0 + 0.5) / 2 = 75\% on this 2-example devset. (The log line prints first because logging and printing go to different streams.)

In an optimization pipeline

This is stage 2, measuring the program, and it steers all of stage 3. dspy.BootstrapFewShot uses the metric with a trace to decide which runs become demos. dspy.MIPROv2 uses the plain score to rank instruction and demo combinations. dspy.GEPA uses the score to keep candidates on its per-example Pareto front and the feedback text to write new instructions. A metric that is wrong in a subtle way is wrong in every one of those places at once.

Common mistakes

  • Score-only feedback for GEPA. A plain number becomes "This trajectory got a score of 0.0." Symptom: GEPA's proposed instructions are vague restatements of the task.
  • Metrics that are easy to game. "Score 1 if the gold answer appears anywhere in the reply" rewards a reply that lists every capital in the world. Symptom: scores climb while replies get longer.
  • Returning a number outside [0, 1]. A 0 to 10 rubric score breaks failure_score=0.0 and perfect_score=1.0 assumptions in GEPA. Divide by the maximum.
  • Crashing on odd output. pred.answer.lower() fails if the answer is None after a parse failure. Guard with (pred.answer or "").
  • Feedback that leaks the answer for every example. Telling the reflection model each gold answer invites instructions that memorize valset quirks. Explain the kind of mistake; quoting the expected value for one example is fine, building a lookup table is not.

Cost

A rule-based metric like the ones above costs microseconds per call, and an optimizer calls it once per rollout, so a GEPA budget of MM metric calls (its max_metric_calls) means MM metric calls and about MM program runs. The cost appears when the metric itself calls a model, as an LLM judge does: then each metric call adds jj judge calls, and a run costs about M⋅(k⋅ctask+j⋅cjudge)M \cdot (k \cdot c_{\text{task}} + j \cdot c_{\text{judge}}) dollars, with kk predictors per program run and cc the average cost of one call of each kind. Builder time is the real expense: a good feedback metric takes an hour or two to write and a few rounds of reading GEPA's proposals to tune.

Going further

Leads to

Back to DSPy and GEPA: programming and optimizing language-model systems