prerequisite
Training, validation, and test splits
Why you split examples into sets with different jobs, and how using the wrong set gives a score that lies.
Before this
Nothing beyond first-year college math. This is a starting page.
Why you need this
Every optimizer in this cluster makes choices by scoring candidates on examples. If the examples it chooses with are the same ones you report on, the reported score is too high, sometimes by a lot. Splitting your examples into sets with separate jobs is the cheapest protection there is, and it has to be decided before the first optimization run, not after.
The idea
You want a program that works on inputs it has never seen. That property is called generalization. The danger is overfitting: doing well on the examples you tuned on because of quirks of those examples, not because the program got better at the task.
A plain-language picture: a student who memorizes last year's exam answers scores 100% on last year's exam and badly on this year's. Last year's exam was fine for practice; it just cannot tell you how the student will do on a new one.
So you split your labeled examples into three sets, each with one job:
| Set | Name in DSPy | Job | May the optimizer look at it? |
|---|---|---|---|
| Training set | trainset |
Learn from: worked examples are drawn from it, and GEPA reflects on runs over it | Yes, as much as it likes |
| Validation set | valset |
Choose between: score candidates and keep the best | Yes, but only to score |
| Test set | testset |
Report: one final, honest number | No. Used once, at the end |
(Some DSPy functions call the scoring set devset, short for development set; dspy.Evaluate(devset=...) is one. It is the same idea as a validation set.)
Why choosing on the test set leaks
Choosing is a form of learning. Suppose you have 10 candidate instructions that are all, in truth, equally good: each answers 70% of all possible inputs correctly. You score each on the same 20 examples and keep the highest. Each score wobbles by chance, and you kept the luckiest one. A seeded simulation of 100,000 such contests (checked with a short node script) gives:
| Candidates compared | Expected score of the winner on its 20 examples | True accuracy of the winner |
|---|---|---|
| 1 | 70.0% | 70% |
| 5 | 81.5% | 70% |
| 10 | 85.0% | 70% |
Nothing improved, yet the reported number rose 15 points. If those 20 examples were your test set, you would publish 85% for a 70% program. The test set exists so that the final number comes from examples nobody chose with. An optimizer like GEPA compares far more than 10 candidates, so this effect is the normal case, not a corner case.
Small-data reality
DSPy is built to work with small datasets: tens to a few hundred examples, not the tens of thousands that training a model from scratch needs. The dspy.GEPA source in DSPy 3.4.0 even logs a suggestion to use a smaller valset once it has more than 35 examples, to save budget for exploring. Small sets make the split matter more, because every example you move to one set is taken from another. Noisy scores and sample size works out how much a score on 20 examples wobbles.
Stratifying
If your examples have labels, a random split can leave one set short of a label. Stratifying means splitting each label group separately in the same proportions, so every set has the same mix as the whole.
Worked example
You have 60 labeled product reviews for a sentiment task: 30 positive, 18 negative, 12 neutral. You split them 20 / 20 / 20.
Step 1: stratify. Split each label into thirds.
| Label | Total | trainset | valset | testset | Share of each set |
|---|---|---|---|---|---|
| positive | 30 | 10 | 10 | 10 | 50% |
| negative | 18 | 6 | 6 | 6 | 30% |
| neutral | 12 | 4 | 4 | 4 | 20% |
| all | 60 | 20 | 20 | 20 | 100% |
Without stratifying, a random 20 of the 60 holds 2 or fewer neutral reviews about 15% of the time (0.152, from the hypergeometric formula, checked with node). A valset with 1 neutral review cannot tell you whether the program handles neutral reviews at all.
Step 2: shuffle within each label with a fixed seed, so the split is the same every time you rerun it, and write the three lists to files. From now on they do not change.
Step 3: give each optimizer step only the sets it is allowed.
| Step | trainset (20) | valset (20) | testset (20) |
|---|---|---|---|
| Baseline score of the unoptimized program | scores | ||
dspy.LabeledFewShot picks demos (worked examples) |
draws demos | ||
dspy.BootstrapFewShot runs the teacher and keeps passing traces as demos |
runs and keeps | ||
dspy.MIPROv2 bootstraps demos and proposes instructions |
runs and proposes | ||
dspy.MIPROv2 scores each combination it tries |
scores | ||
dspy.GEPA minibatches for reflection (3 examples at a time by default) |
runs, reflects | ||
dspy.GEPA per-example scores for its Pareto front of candidates |
scores | ||
| You pick the final program among optimizer runs | scores | ||
| Final reported number | scored once |
Read the columns: the testset column has exactly one entry. If you look at the test score and then change anything (a new run, a different budget, a tweaked metric), the test set has become a second valset, and you need fresh test examples for an honest number.
In an optimization pipeline
Splits belong to stage 2, measuring the program, and every later stage depends on them. Stage 3 optimizers take trainset and valset directly; compile(program, trainset=..., valset=...) is the usual call. Watch the defaults: in the 3.4.0 source, dspy.GEPA with no valset uses the trainset for both jobs, and dspy.MIPROv2 with no valset carves one out of the trainset (80% of it). Pass your own. In stage 5, the testset is how you notice that an optimized program overfit its valset.
Common mistakes
- Reporting the valset score. The optimizer chose its winner on the valset, so that score is inflated. Symptom: the program does worse in real use than the number you shared.
- Splitting after looking. Writing hard examples, scoring the program, then moving the failures into the trainset to "help it". The test set then lacks the hard cases. Symptom: a high test score and complaints from users.
- Duplicates across sets. The same question worded twice, once in train and once in test. The program can copy a demo's answer. Symptom: suspiciously perfect scores on a few test examples.
- Re-splitting every run. A new random split each time makes scores from different runs incomparable. Fix the seed and save the lists.
- A testset too small to say anything. With 5 test examples, one answer moves the score 20 points.
Cost
Splitting itself is free: shuffling and slicing examples takes time proportional to , which is nothing. The real cost is examples. Each example held out for validation or testing is one the optimizer cannot learn from, and labeled examples cost the builder time to write and check. Scoring cost follows the sets: an optimizer that scores candidates on a valset of examples with a -predictor program makes about model calls for scoring alone, so a valset twice as large doubles that part of the bill. The test set is the cheapest of the three to run, calls for test examples, because it runs once.
Going further
- Noisy scores and sample size: how big each set should be for a difference to be real.
- Building a dataset: the same idea in DSPy code, with
dspy.Exampleand a seeded stratified split. - Read about k-fold cross-validation, the standard way to reuse a small dataset for both training and validation, and why it is rarely used with prompt optimizers (every fold multiplies model calls).
Leads to
- techniqueBuilding a datasetTurning examples into dspy.Example objects, marking which fields are inputs, deciding how many you need, and splitting them for optimization.
- prerequisiteNoisy scores and sample sizeWhy a score measured on a small set wobbles, how to estimate the wobble, and how many examples you need before a difference is real.
Back to DSPy and GEPA: programming and optimizing language-model systems