technique
Few-shot bootstrapping
The first family of optimizers: LabeledFewShot, BootstrapFewShot, and random search over bootstrapped demos, which pick worked examples for every predictor automatically.
Before this
This page assumes you are comfortable with:
- techniqueEvaluating programsRunning a program over a dataset with dspy.Evaluate, reading the results, and deciding whether a change is a real improvement or noise.
- techniqueComposing programsBuilding a multi-step program as a dspy.Module subclass: several predictors wired together in forward, with tools, loops, and branches in plain Python.
Why you need this
A program that works zero-shot on your laptop often gets better the moment its prompt carries a few worked examples. Picking those examples by hand is slow, and in a multi-step program you would also have to write examples for the middle steps, which nobody labeled. Few-shot bootstrapping is stage 3 of the pipeline at its cheapest: an optimizer runs a strong model through your program, keeps the runs your metric accepts, and turns them into examples for every step.
The idea
A program is a DSPy module tree. Each predictor in it (one dspy.Predict call) has two things an optimizer can change: its instruction and its list of demos, the worked examples placed in its prompt. This family of optimizers changes only the demos.
Every optimizer in DSPy has the same entry point:
compiled = optimizer.compile(program, trainset=trainset)
compile copies your program, fills in the copy's demos, and returns it. Your original is untouched. The trainset is a list of dspy.Example objects with their input fields marked.
LabeledFewShot
dspy.LabeledFewShot(k=16) makes no model calls at all. It samples up to k examples from the trainset (with a fixed random seed) and attaches them to every predictor as demos. It is a useful baseline and costs nothing. Its weakness: a training example only carries the fields you labeled. In a two-step program where step one writes a summary and step two classifies it, your examples have a ticket and a category but no summary, so they fit neither step.
BootstrapFewShot
dspy.BootstrapFewShot fixes that by generating the missing fields. In plain words: run a teacher program on each training example, record what every predictor saw and said, score the final answer with your metric, and keep the whole record only if the metric passes. The record is a trace: the inputs and outputs of every predictor during one rollout (one run of the program on one example).
Precisely, from the DSPy 3.4 source:
- The teacher defaults to a copy of your program. If
max_labeled_demosis above zero, the teacher first gets labeled demos fromLabeledFewShot, with the current example removed so it cannot copy the answer. - For each training example in order, the teacher runs. If
metric(example, prediction, trace)is truthy (or at leastmetric_threshold, when you set one), each step of the trace becomes a demo for the predictor that produced it, markedaugmented=True. - It stops once
max_bootstrapped_demosexamples have passed (default 4). Withmax_roundsabove 1, a failed example is retried at temperature 1.0 to get a different answer. - Each predictor in the returned program gets its bootstrapped demos first, then labeled examples that were not bootstrapped, until it has
max_labeled_demosin total (default 16).
The teacher is usually the same program run on a stronger model. You pass that model through teacher_settings, which DSPy applies as a context while the teacher runs. In this cluster the teacher is Claude Opus 5.5 and the student runs on the local task model served by Ollama:
# Excerpt; needs a real model. The key comes from ANTHROPIC_API_KEY.
teacher = dspy.LM("anthropic/claude-opus-5-5", temperature=1.0, max_tokens=32000)
optimizer = dspy.BootstrapFewShot(metric=metric, teacher_settings=dict(lm=teacher))
BootstrapFewShotWithRandomSearch
One bootstrap run gives one set of demos, and the first few passing examples are not necessarily the best ones. dspy.BootstrapFewShotWithRandomSearch builds many candidate programs and keeps the best on a validation set. With num_candidate_programs=16 (the default) it tries 19 candidates: zero-shot, labeled-only, one unshuffled bootstrap, then 16 bootstraps on shuffled trainsets, each with a random number of bootstrapped demos between 1 and max_bootstrapped_demos. Each candidate is scored with dspy.Evaluate on the valset (or the trainset if you pass none), and the best-scoring program is returned with the others attached, sorted, as candidate_programs.
What this family cannot do
None of these optimizers touches an instruction. If the instruction says the wrong thing, or the task needs a rule that examples do not show, demos will not fix it. MIPROv2 and GEPA rewrite instructions.
Worked example
A two-step support-ticket program: summarize turns a ticket into a one-line summary, classify turns the summary into billing, bug, or account. The trainset has five tickets labeled only with a category. Two DummyLM objects stand in for the models, so no network call is made: the teacher's answers are scripted per input, and the teacher gets one ticket wrong on purpose (it calls a refund request account).
This is the complete file, run with DSPy 3.4.0 on Python 3.14 (3.12 and newer behave the same):
"""BootstrapFewShot on a two-step ticket triage program, run on DummyLM."""
import dspy
from dspy.utils import DummyLM
class Triage(dspy.Module):
def __init__(self):
super().__init__()
self.summarize = dspy.Predict("ticket -> summary")
self.classify = dspy.Predict("summary -> category")
def forward(self, ticket):
summary = self.summarize(ticket=ticket).summary
return self.classify(summary=summary)
def ex(ticket, category):
return dspy.Example(ticket=ticket, category=category).with_inputs("ticket")
trainset = [
ex("I was charged twice for my March invoice.", "billing"),
ex("Can I get a refund for the months I did not use?", "billing"),
ex("The app crashes when I open Settings.", "bug"),
ex("How do I change the email on my account?", "account"),
ex("Export to CSV drops the last row.", "bug"),
]
def metric(example, pred, trace=None):
return example.category == pred.category
# Stands in for the teacher (Claude Opus 5.5). Keys are matched against the
# last message of each prompt, so every input gets a fixed scripted answer.
teacher_lm = DummyLM({
"charged twice": {"summary": "Duplicate charge on the March invoice."},
"Duplicate charge": {"category": "billing"},
"refund for the months": {"summary": "Customer wants money back for unused months."},
"money back": {"category": "account"}, # wrong: the gold label is billing
"crashes when I open": {"summary": "App crashes on opening Settings."},
"App crashes": {"category": "bug"},
})
# Stands in for the local task model.
student_lm = DummyLM({
"password reset email": {"summary": "Password reset email never arrives."},
"Password reset": {"category": "account"},
})
dspy.configure(lm=student_lm)
optimizer = dspy.BootstrapFewShot(
metric=metric,
max_bootstrapped_demos=2,
max_labeled_demos=3,
teacher_settings=dict(lm=teacher_lm),
)
compiled = optimizer.compile(Triage(), trainset=trainset)
for name, predictor in compiled.named_predictors():
print(f"\n{name}: {len(predictor.demos)} demos")
for demo in predictor.demos:
kind = "bootstrapped" if demo.get("augmented") else "labeled"
fields = {k: v for k, v in demo.items() if k != "augmented"}
print(f" [{kind}] {fields}")
pred = compiled(ticket="The password reset email never shows up.")
print("\nprediction:", pred.category)
dspy.inspect_history(n=1)
Run it with python bootstrap_triage.py. The output, with the progress bar and the system message trimmed:
Bootstrapped 2 full traces after 3 examples for up to 1 rounds, amounting to 3 attempts.
summarize: 3 demos
[bootstrapped] {'ticket': 'I was charged twice for my March invoice.', 'summary': 'Duplicate charge on the March invoice.'}
[bootstrapped] {'ticket': 'The app crashes when I open Settings.', 'summary': 'App crashes on opening Settings.'}
[labeled] {'ticket': 'Export to CSV drops the last row.', 'category': 'bug'}
classify: 3 demos
[bootstrapped] {'summary': 'Duplicate charge on the March invoice.', 'category': 'billing'}
[bootstrapped] {'summary': 'App crashes on opening Settings.', 'category': 'bug'}
[labeled] {'ticket': 'Export to CSV drops the last row.', 'category': 'bug'}
prediction: account
...
User message:
[[ ## summary ## ]]
Duplicate charge on the March invoice.
Assistant message:
[[ ## category ## ]]
billing
[[ ## completed ## ]]
User message:
[[ ## summary ## ]]
App crashes on opening Settings.
Assistant message:
[[ ## category ## ]]
bug
[[ ## completed ## ]]
User message:
[[ ## summary ## ]]
Password reset email never arrives.
...
Read it line by line:
| Trainset example | Teacher's final answer | Metric | Result |
|---|---|---|---|
| 1, double charge | billing | pass | trace kept: one demo for each predictor |
| 2, refund | account | fail | trace thrown away |
| 3, crash | bug | pass | trace kept; 2 reached, so it stops |
| 4, email change | not run | left as a labeled example | |
| 5, CSV export | not run | left as a labeled example |
Three points are worth seeing in the output:
- The
summarizedemos contain summaries that no person wrote. The teacher produced them, and the metric on the final category vouched for them. That is how bootstrapping labels intermediate steps. - Each predictor got one labeled example to fill its third slot (sampled from examples 2, 4, and 5, the ones not bootstrapped). But that example has a
ticketand acategory, and the printed prompt does not contain it: the adapter skipped it because it lacks thesummaryfield. In a multi-step program, labeled demos only help a predictor whose fields they actually contain. - The wrong teacher answer never reached a prompt. The metric is the filter, so a metric that accepts wrong answers lets wrong demos through.
In an optimization pipeline
This is the first optimizer to try at stage 3, right after stage 2 gave you a baseline score you trust (Evaluating programs). Run LabeledFewShot and BootstrapFewShot, score both on the valset, and compare them with the zero-shot baseline. If bootstrapping helps and you have budget, random search usually adds a little more. If instructions are the problem, move on to MIPROv2 (which reuses this exact bootstrapping as its first step) or GEPA. Bootstrapped traces also become training data in stage 4: BootstrapFinetune is the same filtering idea with fine-tuning at the end.
Common mistakes
- A loose metric. If the metric returns true for half-right answers, bad traces become demos. Symptom: the compiled program confidently repeats a pattern from a wrong demo, for example labeling every refund
account. - Using the same model as teacher and student, then expecting a big jump. The demos can only be as good as the teacher's passing runs. Symptom: compiled score within noise of the zero-shot score.
- Forgetting that labeled examples may not fit. In a multi-step program,
max_labeled_demosslots filled with examples missing a predictor's fields are silently dropped from the prompt. Symptom:len(predictor.demos)says 3, the prompt shows 2. - Choosing among random-search candidates on the trainset. Without a
valset,BootstrapFewShotWithRandomSearchscores candidates on the same examples the demos came from. Symptom: a high compile-time score that falls back on the testset. - Too many demos for the context window. Sixteen long demos per predictor can crowd out the input. Symptom: truncated inputs or rising latency per call.
Cost
Let be the trainset size, the number of predictors, the value of max_rounds, and the valset size. LabeledFewShot makes no model calls. BootstrapFewShot runs at most teacher rollouts of calls each, and usually far fewer, because it stops when max_bootstrapped_demos traces have passed; the example above stopped after 3 rollouts, which is 6 teacher calls. With an expensive teacher these are the only calls that cost dollars: teacher dollars , where is the number of teacher calls, is tokens per call, and is the provider's price per token. BootstrapFewShotWithRandomSearch adds one full evaluation per candidate: with the default 16 candidates plus the 3 fixed ones and a 50-example valset, that is student rollouts, plus up to 19 bootstrap runs. The cost that lasts is in serving: every demo is sent with every call forever after, so a predictor with demos of tokens each pays about extra input tokens per call.
Going further
- MIPROv2, which bootstraps demo sets this way and then also proposes instructions.
- GEPA: reflective prompt evolution, for when instructions, not examples, are the bottleneck.
- BootstrapFinetune, which trains on the traces instead of placing them in the prompt.
- The DSPy documentation page on the BootstrapFewShot family, and the source file
dspy/teleprompt/bootstrap.py, which is short enough to read in one sitting. dspy.KNNFewShot, which picks demos per input by similarity instead of fixing one set.
Leads to
- techniqueBootstrapFinetuneDistilling a program into a smaller model's weights: collecting traces that pass the metric with a strong teacher, then fine-tuning the local task model on them.
- techniqueMIPROv2Optimizing instructions and examples together: proposing candidate instructions, bootstrapping demo sets, and searching combinations with Bayesian optimization.
Back to DSPy and GEPA: programming and optimizing language-model systems