prerequisite

Training, validation, and test splits

Why you split examples into sets with different jobs, and how using the wrong set gives a score that lies.

Before this

Nothing beyond first-year college math. This is a starting page.

Why you need this

Every optimizer in this cluster makes choices by scoring candidates on examples. If the examples it chooses with are the same ones you report on, the reported score is too high, sometimes by a lot. Splitting your examples into sets with separate jobs is the cheapest protection there is, and it has to be decided before the first optimization run, not after.

The idea

You want a program that works on inputs it has never seen. That property is called generalization. The danger is overfitting: doing well on the examples you tuned on because of quirks of those examples, not because the program got better at the task.

A plain-language picture: a student who memorizes last year's exam answers scores 100% on last year's exam and badly on this year's. Last year's exam was fine for practice; it just cannot tell you how the student will do on a new one.

So you split your labeled examples into three sets, each with one job:

Set Name in DSPy Job May the optimizer look at it?
Training set trainset Learn from: worked examples are drawn from it, and GEPA reflects on runs over it Yes, as much as it likes
Validation set valset Choose between: score candidates and keep the best Yes, but only to score
Test set testset Report: one final, honest number No. Used once, at the end

(Some DSPy functions call the scoring set devset, short for development set; dspy.Evaluate(devset=...) is one. It is the same idea as a validation set.)

Why choosing on the test set leaks

Choosing is a form of learning. Suppose you have 10 candidate instructions that are all, in truth, equally good: each answers 70% of all possible inputs correctly. You score each on the same 20 examples and keep the highest. Each score wobbles by chance, and you kept the luckiest one. A seeded simulation of 100,000 such contests (checked with a short node script) gives:

Candidates compared Expected score of the winner on its 20 examples True accuracy of the winner
1 70.0% 70%
5 81.5% 70%
10 85.0% 70%

Nothing improved, yet the reported number rose 15 points. If those 20 examples were your test set, you would publish 85% for a 70% program. The test set exists so that the final number comes from examples nobody chose with. An optimizer like GEPA compares far more than 10 candidates, so this effect is the normal case, not a corner case.

Small-data reality

DSPy is built to work with small datasets: tens to a few hundred examples, not the tens of thousands that training a model from scratch needs. The dspy.GEPA source in DSPy 3.4.0 even logs a suggestion to use a smaller valset once it has more than 35 examples, to save budget for exploring. Small sets make the split matter more, because every example you move to one set is taken from another. Noisy scores and sample size works out how much a score on 20 examples wobbles.

Stratifying

If your examples have labels, a random split can leave one set short of a label. Stratifying means splitting each label group separately in the same proportions, so every set has the same mix as the whole.

Worked example

You have 60 labeled product reviews for a sentiment task: 30 positive, 18 negative, 12 neutral. You split them 20 / 20 / 20.

Step 1: stratify. Split each label into thirds.

Label Total trainset valset testset Share of each set
positive 30 10 10 10 50%
negative 18 6 6 6 30%
neutral 12 4 4 4 20%
all 60 20 20 20 100%

Without stratifying, a random 20 of the 60 holds 2 or fewer neutral reviews about 15% of the time (0.152, from the hypergeometric formula, checked with node). A valset with 1 neutral review cannot tell you whether the program handles neutral reviews at all.

Step 2: shuffle within each label with a fixed seed, so the split is the same every time you rerun it, and write the three lists to files. From now on they do not change.

Step 3: give each optimizer step only the sets it is allowed.

Step trainset (20) valset (20) testset (20)
Baseline score of the unoptimized program scores
dspy.LabeledFewShot picks demos (worked examples) draws demos
dspy.BootstrapFewShot runs the teacher and keeps passing traces as demos runs and keeps
dspy.MIPROv2 bootstraps demos and proposes instructions runs and proposes
dspy.MIPROv2 scores each combination it tries scores
dspy.GEPA minibatches for reflection (3 examples at a time by default) runs, reflects
dspy.GEPA per-example scores for its Pareto front of candidates scores
You pick the final program among optimizer runs scores
Final reported number scored once

Read the columns: the testset column has exactly one entry. If you look at the test score and then change anything (a new run, a different budget, a tweaked metric), the test set has become a second valset, and you need fresh test examples for an honest number.

In an optimization pipeline

Splits belong to stage 2, measuring the program, and every later stage depends on them. Stage 3 optimizers take trainset and valset directly; compile(program, trainset=..., valset=...) is the usual call. Watch the defaults: in the 3.4.0 source, dspy.GEPA with no valset uses the trainset for both jobs, and dspy.MIPROv2 with no valset carves one out of the trainset (80% of it). Pass your own. In stage 5, the testset is how you notice that an optimized program overfit its valset.

Common mistakes

  • Reporting the valset score. The optimizer chose its winner on the valset, so that score is inflated. Symptom: the program does worse in real use than the number you shared.
  • Splitting after looking. Writing hard examples, scoring the program, then moving the failures into the trainset to "help it". The test set then lacks the hard cases. Symptom: a high test score and complaints from users.
  • Duplicates across sets. The same question worded twice, once in train and once in test. The program can copy a demo's answer. Symptom: suspiciously perfect scores on a few test examples.
  • Re-splitting every run. A new random split each time makes scores from different runs incomparable. Fix the seed and save the lists.
  • A testset too small to say anything. With 5 test examples, one answer moves the score 20 points.

Cost

Splitting itself is free: shuffling and slicing nn examples takes time proportional to nn, which is nothing. The real cost is examples. Each example held out for validation or testing is one the optimizer cannot learn from, and labeled examples cost the builder time to write and check. Scoring cost follows the sets: an optimizer that scores CC candidates on a valset of vv examples with a kk-predictor program makes about C⋅v⋅kC \cdot v \cdot k model calls for scoring alone, so a valset twice as large doubles that part of the bill. The test set is the cheapest of the three to run, t⋅kt \cdot k calls for tt test examples, because it runs once.

Going further

  • Noisy scores and sample size: how big each set should be for a difference to be real.
  • Building a dataset: the same idea in DSPy code, with dspy.Example and a seeded stratified split.
  • Read about k-fold cross-validation, the standard way to reuse a small dataset for both training and validation, and why it is rarely used with prompt optimizers (every fold multiplies model calls).

Leads to

Back to DSPy and GEPA: programming and optimizing language-model systems