technique
LLM-as-judge metrics
When there is no single right answer: using a model with a rubric as the metric, writing the judge as a DSPy program, and keeping it honest.
Before this
This page assumes you are comfortable with:
Why you need this
A metric is the function an optimizer climbs. For a label or a number, the metric compares the output to a gold answer and you are done. For a summary, an explanation, or an email draft, many different outputs are good and no string comparison can tell them apart. An LLM-as-judge metric hands the grading to a language model with a written rubric. Done carefully, it gives stage 2 of the pipeline a usable score for open-ended tasks, and its written reasons give GEPA the feedback text it learns from.
The idea
In plain words: write down what a good answer must do, as a short list of criteria with points, and ask a strong model to award the points and explain every point it took away.
The precise version has three parts.
- A rubric: criteria , each with a maximum number of points and a sentence saying how to award them.
- A judge: a DSPy program whose signature takes the source, the output, and the rubric, and returns one integer per criterion plus a justification. Because it is a signature, the judge is a predictor like any other, with typed outputs that DSPy parses for you.
- A metric that runs the judge and turns the points into a score in :
where is the score of program on example . With small numbers: a rubric worth points, and a summary that earns of them, scores .
Who should judge? Use the strong model, here Claude Opus 5.5 connected with dspy.LM("anthropic/claude-opus-5-5") and its key read from the ANTHROPIC_API_KEY environment variable. A judge weaker than the program it grades cannot see the program's mistakes. The program being graded runs on the local task model, a small model served by Ollama.
Worked example
The task: summarize a short news item in at most 25 words. The rubric has three criteria.
| Criterion | Points | Award |
|---|---|---|
| coverage | 0 to 2 | 2 if the summary states the main decision and its reason, 1 if only one of them, 0 if neither |
| faithful | 0 to 2 | 2 if every claim is in the source, 1 if one small claim is not, 0 if a claim contradicts the source |
| brief | 0 to 1 | 1 if 25 words or fewer |
The judge's answers below are scripted with DSPy's fake model, DummyLM, keyed by summary text, so the run makes no network call. A real run would set judge_lm to the strong model; nothing else in the file changes. This ran on DSPy 3.4 and Python 3.14; 3.12 and newer behave the same.
import dspy
from dspy.utils import DummyLM
RUBRIC = """coverage (0-2): 2 if the summary states the main decision and its reason, 1 if only one of them, 0 if neither.
faithful (0-2): 2 if every claim is in the source, 1 if one small claim is not, 0 if a claim contradicts the source.
brief (0-1): 1 if the summary is 25 words or fewer, else 0."""
class JudgeSummary(dspy.Signature):
"""Grade the summary of the source text against the rubric. Score each criterion, then justify every point lost."""
source: str = dspy.InputField()
summary: str = dspy.InputField()
rubric: str = dspy.InputField()
coverage: int = dspy.OutputField(desc="0, 1, or 2")
faithful: int = dspy.OutputField(desc="0, 1, or 2")
brief: int = dspy.OutputField(desc="0 or 1")
justification: str = dspy.OutputField(desc="one sentence per point lost; empty if none")
judge = dspy.Predict(JudgeSummary)
source = (
"The library board voted 5 to 2 on Tuesday to close the Elm Street branch in June. "
"Members said repair costs for the 1950s building were higher than the cost of expanding "
"the main branch two miles away. Bus service between the two sites will be added."
)
summaries = [
"The board voted to close the Elm Street branch in June because repairs cost more than expanding the main branch.",
"The board voted to close the Elm Street branch in June after residents complained about parking.",
"The library board met on Tuesday and discussed several topics related to branch buildings, repair budgets, bus routes, and long-term planning for the whole library system.",
]
# Scripted judge answers, keyed by summary text. A real judge (the strong model) would produce these.
judge_lm = DummyLM({
summaries[0]: {"coverage": 2, "faithful": 2, "brief": 1, "justification": ""},
summaries[1]: {"coverage": 1, "faithful": 0, "brief": 1,
"justification": "Coverage: the reason is wrong; the source gives repair costs, not parking. "
"Faithful: parking complaints are not in the source and contradict the stated reason."},
summaries[2]: {"coverage": 0, "faithful": 2, "brief": 0,
"justification": "Coverage: neither the decision to close Elm Street nor the reason is stated. "
"Brief: 26 words, over the 25-word limit."},
})
def judge_metric(gold, pred, trace=None, pred_name=None, pred_trace=None):
with dspy.context(lm=judge_lm):
grade = judge(source=gold.source, summary=pred.summary, rubric=RUBRIC)
points = grade.coverage + grade.faithful + grade.brief
score = points / 5
feedback = f"Rubric points {points}/5. {grade.justification}".strip()
return dspy.Prediction(score=score, feedback=feedback)
gold = dspy.Example(source=source).with_inputs("source")
for i, s in enumerate(summaries, start=1):
result = judge_metric(gold, dspy.Prediction(summary=s))
print(f"summary {i}: score={result.score}")
print(f" feedback: {result.feedback}")
Run it with python judge.py. It printed:
summary 1: score=1.0
feedback: Rubric points 5/5.
summary 2: score=0.4
feedback: Rubric points 2/5. Coverage: the reason is wrong; the source gives repair costs, not parking. Faithful: parking complaints are not in the source and contradict the stated reason.
summary 3: score=0.4
feedback: Rubric points 2/5. Coverage: neither the decision to close Elm Street nor the reason is stated. Brief: 26 words, over the 25-word limit.
Summaries 2 and 3 both score 0.4, for opposite reasons: one invents a fact, the other is vague and long. A score alone cannot tell an optimizer which fix each one needs. The feedback can.
How the justification becomes GEPA feedback. judge_metric has the argument list GEPA's metric expects in DSPy 3.4: (gold, pred, trace, pred_name, pred_trace). It returns dspy.Prediction(score=..., feedback=...), which is the shape dspy.GEPA reads. When GEPA reflects on a failed summary, the reflection model sees "parking complaints are not in the source", not just "0.4", and can propose an instruction such as "state only reasons given in the source". If a metric returns a bare number instead, the DSPy 3.4 source substitutes the sentence "This trajectory got a score of 0.4." as the feedback, which tells the reflection model almost nothing. The same metric also works with dspy.Evaluate, which reads the score and ignores the feedback. See GEPA: reflective prompt evolution for what the reflection model does with it.
Keeping the judge honest
A judge is a model, so it has a model's habits. These are the common ones.
| Problem | What it looks like | Countermeasure |
|---|---|---|
| Position bias | When comparing two outputs, the judge prefers whichever is shown first (or second). | Grade one output at a time against the rubric, or run each pair twice with the order swapped and count only consistent verdicts. |
| Length bias | Longer answers get higher scores even when they say less. | Make length its own criterion with a hard limit, as "brief" does above. |
| Leniency | Scores cluster near the top. | Ask for points lost and why, not an overall impression. |
| Drift | The same summary scores differently after the judge model is updated. | Pin the judge model, and recalibrate whenever it changes. |
| Gaming | The program learns phrasing the judge likes ("This summary is accurate and concise"). | Read high-scoring outputs by hand; keep the judge's rubric out of the program's prompt. |
Calibration means checking the judge against people before trusting it. Grade a small set by hand, run the judge on the same set, and compare. Here is an illustration with eight summaries scored out of 5 (made-up grades, not a real run):
| Summary | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
|---|---|---|---|---|---|---|---|---|
| Human | 5 | 2 | 2 | 4 | 1 | 3 | 5 | 0 |
| Judge | 5 | 2 | 3 | 4 | 3 | 3 | 5 | 1 |
The judge matches exactly on 5 of 8. The average absolute difference is points. Every difference has the judge higher: its mean is 3.25 against the humans' 2.75, so it is lenient by half a point. Summary 5, off by 2, is the one to read: its justification will show which rubric line the judge is applying differently, and that line is what to rewrite.
In an optimization pipeline
The judge belongs to stage 2, but it is called throughout stage 3. During Evaluating programs it is called once per example. During GEPA it is called on every rollout of every minibatch and every valset evaluation, and its justification becomes the feedback the reflection model reads. Building it is a one-time cost; running it is a per-call cost for the whole optimization, which is why the rubric should be settled and calibrated before the optimizer starts.
Common mistakes
- No justification field. Symptom: GEPA's proposals are generic ("be more accurate") because the feedback is only a number.
- An overall 1-to-10 score instead of criteria. Symptom: the same output gets 6, then 8, then 7 on reruns at nonzero temperature, and nobody can say why.
- Judging with the task model. Symptom: the program's own mistakes (an invented fact) earn full faithfulness points, because the judge shares the blind spot.
- Never calibrating. Symptom: the valset score rises steadily while people who read the outputs see no improvement.
- Changing the rubric mid-optimization. Symptom: scores before and after the change are not comparable, and the Pareto front mixes candidates graded by two different rules.
Cost
Every metric call is now a strong-model call. If an evaluation covers examples, a judge call uses input tokens and output tokens, and the judge model charges and per token, the judging alone costs about . The source, the output, and the rubric all count toward , and the justification makes larger than a bare score would. For an optimizer with a budget of metric calls, multiply by instead of : the judge can cost more than the task model it grades. Wall-clock time adds one judge round trip per rollout. The builder's time goes into the rubric and into the hand-graded calibration set, which has to be large enough, and varied enough, to show where the judge and people disagree.
Going further
- GEPA in practice, where a feedback metric like this one drives a real optimization.
- Pairwise preference judging and its order-swap check, as used in chatbot comparison leaderboards.
- Inter-rater agreement statistics such as Cohen's kappa, for comparing judge and human grades more carefully than exact matches.
- Using a smaller judge distilled from the strong one once calibration shows it agrees well enough.
Leads to
Back to DSPy and GEPA: programming and optimizing language-model systems