technique

BootstrapFinetune

Distilling a program into a smaller model's weights: collecting traces that pass the metric with a strong teacher, then fine-tuning the local task model on them.

Before this

This page assumes you are comfortable with:

Why you need this

Prompt optimizers change the words a model reads. At some point the words stop being the bottleneck: the small model you can afford to run thousands of times a day simply is not good enough at the task, however it is asked. BootstrapFinetune is stage 4's first tool. It lets a strong teacher do the task, keeps the runs your metric accepts, and trains the small model's weights on them, so the small model learns to do what the teacher did.

The idea

Fine-tuning means changing a model's weights (the numbers inside it) by showing it input and output pairs and nudging it toward producing those outputs. Few-shot bootstrapping already collects good input and output pairs: it runs a program, checks each run with the metric, and keeps the passing runs as worked examples in the prompt. BootstrapFinetune collects the same kind of runs and uses them as training data instead.

Four terms, defined once for this page:

  • A program is a DSPy module tree; a predictor is one dspy.Predict inside it, one model call.
  • A rollout is one run of the program on one example.
  • A trace is the record of every predictor's inputs and outputs during a rollout.
  • The teacher is the program run with a strong model (here Claude Opus 5.5); the student is the same program run with the small model you want to train.

The steps, as the dspy.BootstrapFinetune source in DSPy 3.4 runs them:

  1. Run the teacher on every trainset example and record each trace and its metric score.
  2. Drop every trace whose score is falsy (False, 0, or 0.0). If no metric is given, keep them all.
  3. For each predictor, turn each of its calls in each kept trace into one chat-format training example: the messages the adapter would send, plus the answer it produced. With multitask=True (the default) every predictor's examples go into one shared dataset per model; with multitask=False each predictor gets its own.
  4. Start one fine-tuning job per (model, dataset) pair, wait for it, and point each student predictor at the newly trained model.

The student keeps its prompt instructions. Its demos stay too, unless you pass exclude_demos=True.

When weights beat prompts

Situation Why weights help
High call volume Every call to a large hosted model costs money; a tuned small model costs only its hardware.
A cheap local model Small models gain the most from training, and you own the weights.
Latency A small model answers faster, and a tuned model needs fewer demos in its prompt, so prompts are shorter.
A stable task Training is slow to redo. It pays off when the task and labels will not change next week.

If none of these hold, stay in stage 3. A good prompt on a strong model is cheaper to build and easier to change.

Worked example

What DSPy 3.4 can actually train

The source decides which models can be fine-tuned. In DSPy 3.4 a model can be trained only if its provider (the object that knows how to launch, train, and stop it) says so. Three providers do:

Provider Where training happens What it needs
OpenAIProvider (chosen automatically for openai/... models) OpenAI's hosted fine-tuning service An OpenAI account and a model that service can fine-tune.
DatabricksProvider A Databricks workspace A Databricks account.
LocalProvider (from dspy.clients.lm_local import LocalProvider) Your machine Hugging Face weights; the torch, transformers, trl, and peft packages for training; SGLang to serve the result.

Ollama, which this cluster uses to run the local task model elsewhere, is not one of them: DSPy 3.4 cannot fine-tune through Ollama. The local path trains Hugging Face weights directly. The LocalProvider source trains on an NVIDIA GPU if CUDA is available, else on Apple's GPU (MPS), else on the CPU; its default settings train for 5 epochs at learning rate η=10−5\eta = 10^{-5} and train every weight unless you pass use_peft=True, which switches to LoRA at rank rLoRA=32r_{\text{LoRA}} = 32. The DSPy classification fine-tuning tutorial says it "requires a local GPU at the moment for inference", calls fine-tuning experimental, and uses meta-llama/Llama-3.2-1B-Instruct as its student.

The code

This sample needs a real model and training hardware; nothing shown was run. It follows the tutorial's shape with Claude Opus 5.5 as the teacher.

import dspy
from dspy.clients.lm_local import LocalProvider

classify = dspy.Predict("review -> sentiment")

teacher = classify.deepcopy()
teacher.set_lm(dspy.LM("anthropic/claude-opus-5-5"))  # key from ANTHROPIC_API_KEY

student = classify.deepcopy()
student.set_lm(dspy.LM("openai/local:meta-llama/Llama-3.2-1B-Instruct",
                       provider=LocalProvider(), max_tokens=2000))

def exact_match(example, prediction, trace=None):
    return example.sentiment == prediction.sentiment

trainset = load_reviews()  # your list of dspy.Example, hundreds of them

optimizer = dspy.BootstrapFinetune(metric=exact_match, num_threads=16)
tuned = optimizer.compile(student, teacher=teacher, trainset=trainset)

tuned.get_lm().launch()   # start an SGLang server for the tuned weights
print(tuned(review="Arrived early and works great.").sentiment)
tuned.get_lm().kill()     # free the GPU

Two details from the source matter here. The teacher must have the same predictor names as the student and must not share predictor objects with it, which is why both are deepcopy()s. And every student predictor needs its model set with set_lm; the global dspy.configure model is not enough.

From traces to training data

Steps 1 to 3 do not train anything, so they can run on DSPy's fake model, DummyLM, which returns scripted answers. This file runs a teacher with three scripted answers over three reviews, then calls the same filter-and-format step BootstrapFinetune runs before training (a private method, used here only to show its output). Verified on DSPy 3.4.0 (ran on Python 3.14; 3.12 and newer behave the same).

import json

import dspy
from dspy.teleprompt.bootstrap_trace import bootstrap_trace_data
from dspy.utils import DummyLM

# The fake teacher: three scripted answers, one per review, in order.
teacher_lm = DummyLM([
    {"sentiment": "positive"},
    {"sentiment": "negative"},
    {"sentiment": "positive"},
])

trainset = [
    dspy.Example(review="Arrived early and works great.", sentiment="positive").with_inputs("review"),
    dspy.Example(review="Broke after two days.", sentiment="negative").with_inputs("review"),
    dspy.Example(review="Does what the box says.", sentiment="neutral").with_inputs("review"),
]


def exact_match(example, prediction, trace=None):
    return example.sentiment == prediction.sentiment


teacher = dspy.Predict("review -> sentiment")
teacher.set_lm(teacher_lm)

# Step 1: run the teacher on every example and record the trace and score.
traces = bootstrap_trace_data(program=teacher, dataset=trainset, metric=exact_match, num_threads=1)
for t in traces:
    print(t["example_ind"], repr(t["prediction"].sentiment), "score:", t["score"])

# Step 2: the same filter-and-format step BootstrapFinetune runs before training.
finetuner = dspy.BootstrapFinetune(metric=exact_match)
data, data_format = finetuner._prepare_finetune_data(trace_data=traces, lm=teacher_lm)
print("format:", data_format.value, "| training examples:", len(data))
print(json.dumps(data[0], indent=2))

Run with python traces_to_data.py. Output, with the progress bar trimmed:

... INFO dspy.teleprompt.bootstrap_finetune: Collected data for 3 examples
... INFO dspy.teleprompt.bootstrap_finetune: After filtering with the metric, 2 examples remain

0 'positive' score: True
1 'negative' score: True
2 'positive' score: False
format: chat | training examples: 2
{
  "messages": [
    {
      "role": "system",
      "content": "Your input fields are:\n1. `review` (str):\nYour output fields are:\n1. `sentiment` (str):\n ... Given the fields `review`, produce the fields `sentiment`."
    },
    {
      "role": "user",
      "content": "[[ ## review ## ]]\nArrived early and works great.\n\nRespond with the corresponding output fields, starting with the field `[[ ## sentiment ## ]]`, and then ending with the marker for `[[ ## completed ## ]]`."
    },
    {
      "role": "assistant",
      "content": "[[ ## sentiment ## ]]\npositive\n\n[[ ## completed ## ]]\n"
    }
  ]
}

The same three traces as a table:

Trace Teacher input Teacher output Gold label Score Becomes
0 "Arrived early and works great." positive positive True One chat example: system message with the field layout, user message with the review, assistant message [[ ## sentiment ## ]] positive
1 "Broke after two days." negative negative True One chat example, same shape, assistant says negative
2 "Does what the box says." positive neutral False Dropped: nothing to learn from a wrong answer

The training example is the whole conversation, including the ChatAdapter's field markers. The student learns the format along with the answer, which is why the tuned model should be called through the same adapter it was trained with.

In an optimization pipeline

BootstrapFinetune sits at stage 4 and takes its inputs from the stages before it. Stage 1 gives the program and the two models. Stage 2 gives the metric that filters traces, and a valset and testset to check the tuned student against the teacher. Stage 3 usually runs first: an optimized prompt makes the teacher's traces better, and a better teacher makes better training data. Prompts and weights together turns that ordering into one optimizer. Reinforcement learning with GRPO is the other weight optimizer, trained on scores instead of on kept traces. Stage 5 saves the program, which now names a tuned model that must be kept alongside it.

Common mistakes

  • A partial-credit metric that never filters. The source keeps every trace whose score is truthy, so a score of 0.3 counts as a pass. You see "After filtering with the metric, N examples remain" with N equal to everything collected. Return a bool, or score >= threshold, from the metric you hand to BootstrapFinetune.
  • No model on the student's predictors. Compile stops with "Predictor 0 does not have an LM assigned". Call student.set_lm(...); dspy.configure is not enough.
  • Too few threads. BootstrapFinetune raises a ValueError if num_threads is smaller than the number of training jobs. With multitask=False that is one job per predictor.
  • Expecting Ollama to train. The error says the provider "does not support fine-tuning". Train with LocalProvider on Hugging Face weights, or a hosted provider.
  • Judging the student on the trainset. It was trained on those exact answers. Compare student and teacher on the testset.
  • Too little data. With twenty kept traces the student memorizes them. Collect more trainset examples before training; the DSPy classification tutorial uses a 500-example trainset.

Cost

Let nn be the trainset size and kk the number of predictor calls per rollout. Collecting data costs nn teacher rollouts, so nknk calls to the strong model, at cTc_T dollars per call: Cdata=nkcTC_{\text{data}} = n k c_T. Training costs the hardware for hh hours at cHc_H dollars per hour, or a hosted provider's training fee. After that, each call to the student costs cSc_S instead of the cBc_B the strong model would cost. If the program runs VV times a day, training pays for itself after about

D=nkcT+hcHVk(cB−cS)D = \frac{n k c_T + h c_H}{V k (c_B - c_S)}

days. With n=500n = 500 and k=2k = 2, data collection is 10001000 teacher calls, once. Training itself takes from minutes to hours on a GPU with enough memory to hold the model, its gradients, and its optimizer state; LoRA shrinks that, as LoRA and parameter-efficient tuning explains. The builder's time is the larger cost: setting up the training stack, and redoing the run whenever the task changes.

Going further

  • The DSPy classification fine-tuning tutorial, which runs this whole loop end to end.
  • The BootstrapFinetune and LocalProvider source in DSPy, for the exact training defaults.
  • Reinforcement learning with GRPO, for training on scores rather than kept traces.
  • Prompts and weights together, for combining this with GEPA.

Leads to

Back to DSPy and GEPA: programming and optimizing language-model systems