prerequisite
Noisy scores and sample size
Why a score measured on a small set wobbles, how to estimate the wobble, and how many examples you need before a difference is real.
Before this
This page assumes you are comfortable with:
Why you need this
An optimizer reports that the new program scores 75% and the old one 70%. Is the new one better, or did it get luckier examples? On 20 examples you cannot tell, and an optimizer that keeps the higher number will happily keep a lucky draw. This page gives you a one-line formula for how much a score wobbles and a rule for when a difference is worth believing.
The idea
Run a program on examples. Write the score on example as , with for right and for wrong. The program's score is the average, , which for 0/1 scores is just the proportion of examples it got right.
Think of all the inputs the program will ever see. On those it has some true accuracy , a number you never get to see directly. Your examples are one sample from that world, so is an estimate of , and a different sample of would give a slightly different estimate.
The standard error (SE) is the typical size of that wobble: how far usually lands from . For 0/1 scores it is
In practice you plug your measured in for . Two things to notice: the wobble shrinks with , not , so four times the examples only halves it; and it is largest when is near 0.5.
A rough 95% interval, the range that contains the true about 95 times in 100, is the score plus or minus two standard errors.
The same formula with numbers
At , so (all values checked with a short node script):
| In points | Rough 95% interval | ||
|---|---|---|---|
| 20 | 10.2 | 49.5% to 90.5% | |
| 100 | 4.6 | 60.8% to 79.2% | |
| 400 | 2.3 | 65.4% to 74.6% |
On 20 examples, "70%" means "somewhere from about half to about nine in ten". Getting the SE down to 2.5 points takes examples.
Worked example
Two programs: A is truly right 70% of the time and B 75%. B is better, by 5 points.
Different examples for each
If A and B are each scored on their own random set of examples, the two wobbles add. The SE of the difference is
| SE of A | SE of B | SE of the difference | A scores higher (simulated) | Tie (simulated) | |
|---|---|---|---|---|---|
| 20 | 10.2 | 9.7 | 14.1 | 29.5% | 13.2% |
| 100 | 4.6 | 4.3 | 6.3 | 19.3% | 4.7% |
| 400 | 2.3 | 2.2 | 3.2 | 5.3% | 0.9% |
The last two columns come from a seeded simulation of 100,000 contests per row. At 20 examples, the worse program wins or ties about 43% of the time. Even at 100 examples it wins about one contest in five. The demo below draws the same picture: drag the dev-set size and watch the two histograms pull apart.
The same examples for both: a paired comparison
Usually you score both programs on the same examples, and that helps a lot. Many examples are easy for both or hard for both, so they cancel out. Only the examples where the programs disagree carry information.
Say both programs run on the same 100 examples:
| B right | B wrong | |
|---|---|---|
| A right | 65 | 5 |
| A wrong | 10 | 20 |
A scores out of 100 and B scores . Now look at each example's difference , using the cluster's notation for program 's score on example . It is on 10 examples, on 5, and on the other 85.
- Mean difference: , the same 5 points.
- Spread of the differences: the average of is , so the variance is .
- Standard error: , or 3.8 points.
- Rough 95% interval: , which is to points.
Compare with unpaired scoring at the same : SE 6.3 points, interval to . Pairing cut the wobble by about 40% for free. The answer is still "not proven": the interval includes zero, because the gap of 5 points is only 1.3 standard errors. Keeping this 15% disagreement rate, the gap reaches two standard errors at about examples paired, versus unpaired.
Run-to-run wobble
A third source of noise has nothing to do with which examples you chose. With a sampling temperature above zero, the same program on the same example can answer differently on a rerun: the examples are fixed, but the model's sampling is not. DSPy caches model replies by default, so an exact rerun returns the cached answers and looks perfectly stable. That stability is the cache, not the program. To measure this wobble, rerun with the cache bypassed, as DSPy's own BootstrapFewShot does when it copies the model with a new rollout_id and temperature=1.0.
In an optimization pipeline
This is stage 2, measuring the program. Every optimizer compares candidates by score, so every comparison has this noise in it. dspy.GEPA accepts a new instruction when it beats the old one on a minibatch of 3 examples (the default reflection_minibatch_size), where one example is 33 points, and then confirms on the full valset. MIPROv2 also scores on minibatches before full evaluations. Both are deliberate trades: cheap noisy checks first, expensive precise ones for survivors. Your job is the last comparison, where the same rule applies: compare paired, on the same valset, and doubt any gap under two standard errors.
Common mistakes
- Trusting a gap of a few points on a small valset. On 20 examples one answer is 5 points. Symptom: an "improvement" that vanishes on the test set.
- Comparing scores from different example sets. Program A on last month's valset, program B on this month's. The comparison is unpaired and the sets may differ in difficulty.
- Forgetting the cache. Rerunning to check stability, getting identical numbers, and concluding the program is deterministic.
- Using a bare percentage. "75%" with no cannot be judged. Write "75% on a 100-example valset".
Cost
The formulas cost nothing to compute. Precision costs model calls: halving the standard error takes four times the examples, so four times the calls, tokens, and dollars per evaluation. For a program with predictors evaluated on examples, one evaluation is about calls; with dollars per call on average, that is dollars, paid again for every candidate you score. Labeled examples also cost the builder's time to write. Past a few hundred examples, extra precision rarely changes a decision.
Going further
- Training, validation, and test splits: which set each comparison should use.
- Read about McNemar's test, the textbook test for exactly the paired right/wrong table above.
- Read about the bootstrap (resampling your own examples with replacement), which gives intervals for scores that are not 0/1, such as partial credit.
Leads to
Back to DSPy and GEPA: programming and optimizing language-model systems