prerequisite

Noisy scores and sample size

Why a score measured on a small set wobbles, how to estimate the wobble, and how many examples you need before a difference is real.

Before this

This page assumes you are comfortable with:

Why you need this

An optimizer reports that the new program scores 75% and the old one 70%. Is the new one better, or did it get luckier examples? On 20 examples you cannot tell, and an optimizer that keeps the higher number will happily keep a lucky draw. This page gives you a one-line formula for how much a score wobbles and a rule for when a difference is worth believing.

The idea

Run a program on nn examples. Write the score on example ii as sis_i, with si=1s_i = 1 for right and si=0s_i = 0 for wrong. The program's score is the average, sˉ=1n∑isi\bar{s} = \frac{1}{n}\sum_i s_i, which for 0/1 scores is just the proportion of examples it got right.

Think of all the inputs the program will ever see. On those it has some true accuracy pp, a number you never get to see directly. Your nn examples are one sample from that world, so sˉ\bar{s} is an estimate of pp, and a different sample of nn would give a slightly different estimate.

The standard error (SE) is the typical size of that wobble: how far sˉ\bar{s} usually lands from pp. For 0/1 scores it is

SE=p(1−p)n\text{SE} = \sqrt{\frac{p(1-p)}{n}}

In practice you plug your measured sˉ\bar{s} in for pp. Two things to notice: the wobble shrinks with n\sqrt{n}, not nn, so four times the examples only halves it; and it is largest when pp is near 0.5.

A rough 95% interval, the range that contains the true pp about 95 times in 100, is the score plus or minus two standard errors.

The same formula with numbers

At p=0.7p = 0.7, so p(1−p)=0.21p(1 - p) = 0.21 (all values checked with a short node script):

nn SE=0.21/n\text{SE} = \sqrt{0.21 / n} In points Rough 95% interval
20 0.0105=0.1025\sqrt{0.0105} = 0.1025 10.2 49.5% to 90.5%
100 0.0021=0.0458\sqrt{0.0021} = 0.0458 4.6 60.8% to 79.2%
400 0.000525=0.0229\sqrt{0.000525} = 0.0229 2.3 65.4% to 74.6%

On 20 examples, "70%" means "somewhere from about half to about nine in ten". Getting the SE down to 2.5 points takes n=0.21/0.0252=336n = 0.21 / 0.025^2 = 336 examples.

Worked example

Two programs: A is truly right 70% of the time and B 75%. B is better, by 5 points.

Different examples for each

If A and B are each scored on their own random set of nn examples, the two wobbles add. The SE of the difference is

SEdiff=0.7⋅0.3n+0.75⋅0.25n=0.3975n\text{SE}_{\text{diff}} = \sqrt{\frac{0.7 \cdot 0.3}{n} + \frac{0.75 \cdot 0.25}{n}} = \sqrt{\frac{0.3975}{n}}

nn SE of A SE of B SE of the difference A scores higher (simulated) Tie (simulated)
20 10.2 9.7 14.1 29.5% 13.2%
100 4.6 4.3 6.3 19.3% 4.7%
400 2.3 2.2 3.2 5.3% 0.9%

The last two columns come from a seeded simulation of 100,000 contests per row. At 20 examples, the worse program wins or ties about 43% of the time. Even at 100 examples it wins about one contest in five. The demo below draws the same picture: drag the dev-set size and watch the two histograms pull apart.

The same examples for both: a paired comparison

Usually you score both programs on the same examples, and that helps a lot. Many examples are easy for both or hard for both, so they cancel out. Only the examples where the programs disagree carry information.

Say both programs run on the same 100 examples:

B right B wrong
A right 65 5
A wrong 10 20

A scores 65+5=7065 + 5 = 70 out of 100 and B scores 65+10=7565 + 10 = 75. Now look at each example's difference di=si(B)−si(A)d_i = s_i(B) - s_i(A), using the cluster's notation si(p)s_i(p) for program pp's score on example ii. It is +1+1 on 10 examples, −1-1 on 5, and 00 on the other 85.

  1. Mean difference: dˉ=(10−5)/100=0.05\bar{d} = (10 - 5) / 100 = 0.05, the same 5 points.
  2. Spread of the differences: the average of di2d_i^2 is 15/100=0.1515 / 100 = 0.15, so the variance is 0.15−0.052=0.14750.15 - 0.05^2 = 0.1475.
  3. Standard error: 0.1475/100=0.0384\sqrt{0.1475 / 100} = 0.0384, or 3.8 points.
  4. Rough 95% interval: 0.05±2×0.03840.05 \pm 2 \times 0.0384, which is −2.7-2.7 to 12.712.7 points.

Compare with unpaired scoring at the same n=100n = 100: SE 6.3 points, interval −7.6-7.6 to 17.617.6. Pairing cut the wobble by about 40% for free. The answer is still "not proven": the interval includes zero, because the gap of 5 points is only 1.3 standard errors. Keeping this 15% disagreement rate, the gap reaches two standard errors at about 4×0.1475/0.052=2364 \times 0.1475 / 0.05^2 = 236 examples paired, versus 4×0.3975/0.052=6364 \times 0.3975 / 0.05^2 = 636 unpaired.

Run-to-run wobble

A third source of noise has nothing to do with which examples you chose. With a sampling temperature above zero, the same program on the same example can answer differently on a rerun: the examples are fixed, but the model's sampling is not. DSPy caches model replies by default, so an exact rerun returns the cached answers and looks perfectly stable. That stability is the cache, not the program. To measure this wobble, rerun with the cache bypassed, as DSPy's own BootstrapFewShot does when it copies the model with a new rollout_id and temperature=1.0.

In an optimization pipeline

This is stage 2, measuring the program. Every optimizer compares candidates by score, so every comparison has this noise in it. dspy.GEPA accepts a new instruction when it beats the old one on a minibatch of 3 examples (the default reflection_minibatch_size), where one example is 33 points, and then confirms on the full valset. MIPROv2 also scores on minibatches before full evaluations. Both are deliberate trades: cheap noisy checks first, expensive precise ones for survivors. Your job is the last comparison, where the same rule applies: compare paired, on the same valset, and doubt any gap under two standard errors.

Common mistakes

  • Trusting a gap of a few points on a small valset. On 20 examples one answer is 5 points. Symptom: an "improvement" that vanishes on the test set.
  • Comparing scores from different example sets. Program A on last month's valset, program B on this month's. The comparison is unpaired and the sets may differ in difficulty.
  • Forgetting the cache. Rerunning to check stability, getting identical numbers, and concluding the program is deterministic.
  • Using a bare percentage. "75%" with no nn cannot be judged. Write "75% on a 100-example valset".

Cost

The formulas cost nothing to compute. Precision costs model calls: halving the standard error takes four times the examples, so four times the calls, tokens, and dollars per evaluation. For a program with kk predictors evaluated on nn examples, one evaluation is about k⋅nk \cdot n calls; with cc dollars per call on average, that is knck n c dollars, paid again for every candidate you score. Labeled examples also cost the builder's time to write. Past a few hundred examples, extra precision rarely changes a decision.

Going further

  • Training, validation, and test splits: which set each comparison should use.
  • Read about McNemar's test, the textbook test for exactly the paired right/wrong table above.
  • Read about the bootstrap (resampling your own examples with replacement), which gives intervals for scores that are not 0/1, such as partial credit.

Leads to

Back to DSPy and GEPA: programming and optimizing language-model systems