technique
Overfitting and re-optimizing
Knowing when an optimized program has stopped being good: overfitting to the valset, the model changing underneath you, data drift, and when to run the optimizer again.
Before this
This page assumes you are comfortable with:
- techniqueSaving and serving programsKeeping an optimized program: saving state versus the whole program, loading it in production, pinning models and versions, and serving it behind an API.
- techniqueEvaluating programsRunning a program over a dataset with dspy.Evaluate, reading the results, and deciding whether a change is a real improvement or noise.
Why you need this
An optimizer reports a score, and the score is the number it was trying to raise. That makes it the least trustworthy number you have. After you ship, three things can make a good program bad: it was fitted to the valset rather than the task, the model underneath it changed, or the inputs changed. This page is how to notice each one, and how to decide whether to run the optimizer again.
The idea
Three splits, defined once for this page. The trainset is what the optimizer learns from (demos come from it, GEPA's minibatches are drawn from it). The valset is what the optimizer chooses with: every candidate is scored on it, and the best one wins. The testset is kept out of optimization entirely and used only to report the final number.
Overfitting means the program got better at the examples it was chosen on without getting better at the task. With prompt optimizers it happens in two ways:
- Selection on noise. GEPA or MIPROv2 scores many candidates on the same valset and keeps the best. Some candidates score high partly by luck on those particular examples. The winner's valset score is the maximum of many noisy numbers, so it is biased upward even if no candidate is truly better.
- Memorized quirks. A reflection model reading valset failures can write the failures' specifics into the instruction: "If the question mentions the Main Street branch, the answer is 9 am." That fixes one valset example and helps nothing else, or hurts when the hours change.
The sign is a gap: the valset score rises and the testset score stays flat. That is why the testset is the only honest final number, and why it must not be used to choose anything. Once you pick between programs by testset score, it has become a second valset.
How big a gap is real? For a score (a fraction between 0 and 1) measured on examples, the standard error is , and a rough 95% range is (Noisy scores and sample size derives this). At on a 200-example testset, , so the range is about points. At on a 50-example valset, , about points. Small valsets make overfitting both more likely and harder to see.
Things that change underneath you
| Change | What happens | What to do |
|---|---|---|
| The local task model's tag gets new weights, or you switch models | The optimized instruction and demos were chosen for the old model. They may now be wrong, or merely unnecessary. | Re-evaluate the saved program on the testset with the new model. Re-optimize only if the score dropped beyond the noise. |
| The reflection or teacher model changes (a new Claude Opus version) | The shipped program never calls it, so nothing changes until the next run. Runs before and after are not exactly comparable. | Record the version in the run record. If an LLM judge is your metric, recheck it against your human-graded examples, because the scores themselves move. |
| DSPy is upgraded | The adapter may format the same saved state into different messages. | Re-evaluate before shipping the upgrade; pin the version as Saving and serving programs describes. |
| Data drift: new kinds of input appear | The testset no longer looks like real traffic, so a steady testset score hides a falling real one. | Label a fresh sample of recent inputs and score it. Fold new kinds of example into all three splits. |
Worked example
Three runs over time
This table is an illustration, not a record of real runs. The task is a short-answer bot for a library website's questions (hours, loans, accounts). Splits: a 60-example trainset, a 50-example valset, a 200-example testset. The optimizer is GEPA; scores are percentages.
| Run | Trigger | Before | After | Decision |
|---|---|---|---|---|
| 1, week 0 | First optimization | valset 58%, testset 57% | valset 78%, testset 70% | Ship. The testset gain of 13 points is twice the noise range. |
| 2, week 9 | Task model tag updated | Shipped program re-evaluated: testset 62% (was 70%) | Re-optimized with the new model: valset 80%, testset 71% | Ship. The 8-point drop exceeded the range at 62%, so it was real; the new run recovered it. |
| 3, week 20 | Library launched e-book lending; complaints | 60 recent real questions, labeled: 55% | Heavier GEPA budget on the same splits: valset 90%, testset 70%, recent 57% | Reject. Valset +10, testset flat, recent sample within noise ( at 55% of 60). Add e-book questions to all three splits and run again. |
Run 3 shows both failure modes at once. The valset had now been used to choose winners in three runs, and a bigger budget simply found a candidate that fit it better. Meanwhile the drift was invisible to the testset, because the testset had no e-book questions. The recent sample is what revealed it.
A dated run record
Every optimization run gets one line in a log, kept next to the saved programs. This file appends a record and prints the log. The values are the illustration's run 1; the date is made up. Verified on DSPy 3.4.0 (ran on Python 3.14; 3.12 and newer behave the same) (it makes no model call).
import json
from datetime import date
import dspy
record = {
"date": date(2026, 3, 2).isoformat(),
"dspy": dspy.__version__,
"optimizer": "GEPA",
"budget": {"auto": "light"},
"task_model": "ollama_chat/llama3.2",
"reflection_model": "anthropic/claude-opus-5-5",
"splits": {"trainset": 60, "valset": 50, "testset": 200},
"scores": {"valset_before": 0.58, "valset_after": 0.78, "testset_before": 0.57, "testset_after": 0.70},
"saved_as": "faq_bot_2026-03-02.json",
"decision": "ship",
}
with open("runs.jsonl", "a") as f:
f.write(json.dumps(record) + "\n")
with open("runs.jsonl") as f:
runs = [json.loads(line) for line in f]
for r in runs:
gain_val = r["scores"]["valset_after"] - r["scores"]["valset_before"]
gain_test = r["scores"]["testset_after"] - r["scores"]["testset_before"]
print(r["date"], r["optimizer"], f"valset {gain_val:+.2f}", f"testset {gain_test:+.2f}", r["decision"])
Run with python run_record.py:
2026-03-02 GEPA valset +0.20 testset +0.13 ship
Record the exact tag you ran; llama3.2 is the tag this cluster uses. The two gains side by side are the overfitting check: a large valset gain next to a near-zero testset gain is the warning.
Re-optimization trigger checklist
Re-evaluate the saved program on the testset (and on a fresh sample of recent inputs) when any of these happen. Re-optimize only if the score falls by more than about two standard errors.
- The task model's tag, weights, or provider changed.
- DSPy was upgraded.
- A scheduled re-evaluation (monthly, say) shows a drop beyond the noise range.
- A labeled sample of recent real inputs scores clearly below the testset.
- New kinds of input, or new labels, appeared in the task.
- The metric or the judge model changed. Scores before and after are not comparable: re-score the baseline first.
Not triggers on their own: one bad report from a user, a valset score that "could be higher", or a new optimizer release.
In an optimization pipeline
This is the last step of stage 5, and it loops back. A re-optimization goes to stage 3 (or 4) with the same metric from stage 2, unless drift means stage 2's dataset needs new examples first, which is the usual case. The saved JSON from Saving and serving programs is the baseline every new run must beat on the testset. Evaluating programs supplies the paired comparison: score old and new programs on the same examples, which is more sensitive than comparing two separate scores.
Common mistakes
- Reporting the valset score as the result. The number looks great and real traffic disagrees. Report the testset.
- Choosing between runs on the testset. After a few rounds the testset score is as optimistic as the valset's. Keep a decision rule fixed in advance (ship if the testset gain exceeds two standard errors), or hold back a second, untouched set.
- A 20-example valset. At 70% its range is about points, so the optimizer is choosing on noise. Move examples from trainset to valset; GEPA in particular needs the valset to be large enough to rank candidates.
- Never reading the optimized instruction. Memorized quirks are visible in plain text: names, numbers, or exact phrasings from valset examples. Read the JSON after every run.
- Re-optimizing on every model update without re-evaluating first. Often the old program is still fine, and the new run only adds selection noise.
- No run record. Months later nobody knows which model, budget, or splits produced the shipped program, so a regression cannot be traced.
Cost
Re-evaluating a saved program costs calls to the local task model, for a testset of examples and predictor calls per rollout, plus a metric call per example (at dollars each if the metric is an LLM judge). For the illustration, local calls: minutes of local compute, and cheap enough to run on a schedule. Labeling a recent sample costs a person's time, a few minutes per example. A re-optimization costs a full optimizer run again (GEPA in practice gives the formula), so the checklist exists to make the cheap check first and the expensive run only when the check says so.
Going further
- Noisy scores and sample size, for standard errors and paired comparisons.
- Training, validation, and test splits, for which step may see which set.
- The GEPA paper by Agrawal and colleagues, on how its Pareto-based selection keeps candidates that win different valset examples.
- Reading about "adaptive data analysis" and reusable holdout sets, for why reusing a test set erodes it.
Back to DSPy and GEPA: programming and optimizing language-model systems