technique

MIPROv2

Optimizing instructions and examples together: proposing candidate instructions, bootstrapping demo sets, and searching combinations with Bayesian optimization.

Before this

This page assumes you are comfortable with:

Why you need this

Few-shot bootstrapping picks good worked examples but never rewrites an instruction. When the instruction is vague, or each step of a multi-step program needs its own wording, you want an optimizer that writes instruction candidates too and then finds which instruction goes best with which set of examples. MIPROv2 does that at stage 3 of the pipeline, with a fixed, predictable budget of model calls.

The idea

In plain words: MIPROv2 builds a menu, then searches it. For each predictor (one dspy.Predict inside the program) it prepares a few sets of demos and a few instructions. A candidate is one choice from each predictor's menu. MIPROv2 tries candidates on the valset, learns which choices tend to score well, and steers later tries toward them.

Precisely, dspy.MIPROv2 in DSPy 3.4 runs three steps inside compile.

Step 1: bootstrap demo sets. It calls the same machinery as BootstrapFewShotWithRandomSearch to build num_fewshot_candidates demo sets per predictor: one empty set, one labeled-only set, one unshuffled bootstrap, and the rest bootstrapped from shuffled trainsets with a random number of demos.

Step 2: propose grounded instructions. A prompt model (in this cluster, Claude Opus 5.5) writes instruction candidates. "Grounded" means the proposer is shown real context, not just the task name. From dspy/propose/grounded_proposer.py, it sees:

Input to the proposer Where it comes from
A dataset description The prompt model first summarizes batches of trainset examples
The program's code and a description of it Program-aware mode, on by default
A description of the predictor being improved The prompt model, from the code and an example run
Up to 3 bootstrapped demos Step 1
The current instruction The program
A random prompting tip One of: none, creative, simple, description, high stakes, persona

Each predictor gets num_instruct_candidates instructions, and slot 0 is always replaced by your original instruction, so the starting point stays on the menu.

Step 3: search combinations with Bayesian optimization. The search space is a grid of choices: for each predictor, an instruction index and a demo-set index. Each trial picks one value per choice, builds that program, and scores it. Bayesian optimization, in plain words, means: keep a running guess of which choices tend to produce good scores, and pick the next trial where the guess looks promising or is still uncertain. MIPROv2 uses the Optuna library's tree-structured Parzen estimator for this, which splits past trials into a better group and a worse group and prefers choices that show up more in the better group. No calculus is needed to use it.

Scoring every trial on the full valset would be expensive, so when the valset has more than 50 examples, each trial is scored on a minibatch of 35 examples (minibatch_size). Every 5 trials (minibatch_full_eval_steps), the combination with the best average minibatch score is scored on the full valset. The unoptimized program is scored first as the baseline, and the program returned is the one with the best full-valset score.

auto: light, medium, heavy

You normally set auto and let MIPROv2 size everything. From AUTO_RUN_SETTINGS in the source:

auto nn valset capped at Instruction candidates (with demos / zero-shot)
light 6 100 3 / 6
medium 12 300 6 / 12
heavy 18 1000 9 / 18

nn is the number of demo sets. The number of trials is

N=max⁡(2mlog⁡2n, 1.5 n)N = \max\left(2 m \log_2 n,\ 1.5\, n\right)

rounded down, where mm is the number of predictors, doubled when demos are being optimized (each predictor then has two choices). For one predictor with demos at light: m=2m = 2, so N=max⁡(2⋅2⋅log⁡26, 9)=max⁡(10.34, 9)N = \max(2 \cdot 2 \cdot \log_2 6,\ 9) = \max(10.34,\ 9), which rounds down to 10 trials. If you set auto=None, you must pass num_candidates to the constructor and num_trials to compile yourself.

Zero-shot mode

Pass max_bootstrapped_demos=0 and max_labeled_demos=0 and MIPROv2 optimizes instructions only. It still bootstraps a few demos (3 per set), because the proposer learns from seeing real inputs and outputs, but it throws them away before step 3, so the final prompt has no examples. Use this when context is tight or when inputs are long.

Worked example

The task: label a support ticket billing, bug, or account, with one dspy.Predict("ticket -> category"). A real run needs the local task model and the prompt model. The call looks like this (needs a real model; nothing shown was run):

task = dspy.LM("ollama_chat/llama3.2", api_base="http://localhost:11434", api_key="")
prompt = dspy.LM("anthropic/claude-opus-5-5", temperature=1.0, max_tokens=32000)  # key from ANTHROPIC_API_KEY
dspy.configure(lm=task)

optimizer = dspy.MIPROv2(metric=metric, prompt_model=prompt, task_model=task, auto="light")
optimized = optimizer.compile(dspy.Predict("ticket -> category"), trainset=trainset, valset=valset)

DummyLM can stand in for both models through steps 1 and 2. The fake task model answers each ticket from a fixed table; the fake prompt model returns a scripted instruction keyed on whichever prompting tip the proposer picked. This complete file was run with DSPy 3.4.0 on Python 3.14 (3.12 and newer behave the same):

"""MIPROv2 steps 1 and 2 on DummyLM. Step 3 needs the optional optuna package."""
import logging

import dspy
from dspy.utils import DummyLM

logging.basicConfig(level=logging.ERROR, format="%(message)s")
logging.getLogger("dspy.teleprompt.mipro_optimizer_v2").setLevel(logging.INFO)


def ex(ticket, category):
    return dspy.Example(ticket=ticket, category=category).with_inputs("ticket")


trainset = [
    ex("I was charged twice this month.", "billing"),
    ex("The app freezes on the login screen.", "bug"),
    ex("Please close my account.", "account"),
    ex("My card was declined but I was still billed.", "billing"),
    ex("Charts never finish loading.", "bug"),
    ex("How do I add a second admin?", "account"),
]
valset = [
    ex("Refund my last payment, please.", "billing"),
    ex("Clicking Save does nothing.", "bug"),
    ex("How do I change my username?", "account"),
]


def metric(gold, pred, trace=None):
    return gold.category == pred.category


# Stands in for the local task model: one fixed answer per ticket.
task_lm = DummyLM({
    "charged twice": {"category": "billing"},
    "freezes": {"category": "bug"},
    "close my account": {"category": "account"},
    "declined": {"category": "billing"},
    "Charts never": {"category": "bug"},
    "second admin": {"category": "account"},
    "Refund": {"category": "billing"},
    "Save does nothing": {"category": "bug"},
    "username": {"category": "account"},
})


# Stands in for the prompt model (Claude Opus 5.5). The proposer picks a random
# prompting tip; the fake model keys its scripted instruction on that tip.
def reply(instruction):
    return {
        "observations": "Short customer support tickets, each labeled billing, bug, or account.",
        "summary": "Short support tickets labeled with one of three categories.",
        "program_description": "A single step that maps a ticket to a category.",
        "module_description": "Reads a ticket and outputs its category.",
        "proposed_instruction": instruction,
    }


prompt_lm = DummyLM({
    "persona": reply("You are a support lead. Label the ticket billing, bug, or account."),
    "concise": reply("Label the ticket: billing, bug, or account."),
    "descriptive": reply("Billing covers charges and refunds, bug covers broken features, "
                         "account covers users and settings. Output one label."),
    "creative": reply("Sort the ticket into the right inbox: billing, bug, or account."),
    "high stakes": reply("A misrouted ticket costs a customer a day. Label it billing, bug, or account."),
    "": reply("Classify the support ticket as billing, bug, or account."),
})
dspy.configure(lm=task_lm)

optimizer = dspy.MIPROv2(
    metric=metric,
    prompt_model=prompt_lm,
    task_model=task_lm,
    auto="light",
    max_bootstrapped_demos=2,
    max_labeled_demos=2,
    num_threads=1,
)
try:
    optimizer.compile(dspy.Predict("ticket -> category"), trainset=trainset, valset=valset)
except ImportError as err:
    print("\nStopped at step 3:", err)

Run with python mipro_dummy.py. The output, with timestamps, progress bars, and some warnings trimmed:

RUNNING WITH THE FOLLOWING LIGHT AUTO RUN SETTINGS:
num_trials: 10
minibatch: False
num_fewshot_candidates: 6
num_instruct_candidates: 3
valset size: 3

==> STEP 1: BOOTSTRAP FEWSHOT EXAMPLES <==
...
Bootstrapping N=6 sets of demonstrations...
Bootstrapping set 1/6
Bootstrapping set 2/6
Bootstrapping set 3/6
Bootstrapped 2 full traces after 2 examples for up to 1 rounds, amounting to 2 attempts.
Bootstrapping set 4/6
Bootstrapped 1 full traces after 1 examples for up to 1 rounds, amounting to 1 attempts.
...
==> STEP 2: PROPOSE INSTRUCTION CANDIDATES <==
...
Proposing N=3 instructions...
...
Proposed Instructions for Predictor 0:
0: Given the fields `ticket`, produce the fields `category`.
1: Sort the ticket into the right inbox: billing, bug, or account.
2: A misrouted ticket costs a customer a day. Label it billing, bug, or account.
...
Stopped at step 3: MIPROv2 requires optional dependency 'optuna'. Install it with `pip install dspy[optuna]`.

The printed settings match the table: light, one predictor with demos, so 10 trials, 6 demo sets, and 3 instructions, the first being the original. The valset has only 3 examples, so minibatch is off. The two new instructions came from the "creative" and "high stakes" tips the proposer drew. Step 3 stopped because the shared environment does not have Optuna installed, which is also what a plain pip install dspy gives you.

Step 3 itself needs a real model. The table below is an illustration of what its log looks like on a larger valset, not a real run: each row is a trial, choices are written as (instruction, demo set), and minibatch scores are on 35 examples.

Trial Choice Scored on Score
1 (0, 0), the original program full 100-example valset 61%
2 (1, 3) minibatch 66%
3 (2, 1) minibatch 54%
4 (1, 4) minibatch 71%
... the sampler favors instruction 1 minibatch ...
7 (1, 4), best minibatch average full 100-example valset 68%

In an optimization pipeline

MIPROv2 is the middle option at stage 3. It needs a trainset to bootstrap from and a valset to search on (if you pass no valset, it takes the last 80% of the trainset as valset, which leaves few examples to bootstrap from, so split it yourself). Run it after few-shot bootstrapping has given you a demo-only score to beat. It shines when the task benefits from both demos and a clearer instruction and when the metric is a plain score with no explanation. The GEPA paper by Agrawal and colleagues reports in its abstract that GEPA outperforms MIPROv2 by over 10%, for example +12% accuracy on AIME-2025; GEPA's advantage comes from reading feedback text, which MIPROv2 never sees.

Common mistakes

  • No Optuna installed. Steps 1 and 2 run, spend prompt-model calls, and then the run stops with an ImportError. Install the optional extra before the first real run.
  • A tiny valset. With 20 examples, two candidates a few points apart are within noise, and the search chases luck. Symptom: the winner scores lower on the testset than the original. See Noisy scores and sample size.
  • Setting num_candidates or num_trials while auto is set. The constructor default is auto="light", so passing num_candidates alone raises an error. Pass auto=None with both.
  • A weak prompt model. The proposer's instructions are only as good as the model writing them. Symptom: every proposed instruction is a reworded copy of the original.
  • Reading the minibatch score as the result. Minibatch scores are on 35 examples and noisy; only full-valset scores decide the winner. Symptom: a log line with a high minibatch score that the returned program does not match.

Cost

Let mm be the number of predictors, VV the valset size after capping, NN the number of trials, MM the minibatch size (35), and cc the number of instruction candidates. From the call estimator in the source: the prompt model makes about 10+c⋅m+(m+1)10 + c \cdot m + (m + 1) calls (dataset summaries, one proposal per instruction per predictor, and program descriptions), which is 15 for one predictor at light. The task model makes V⋅NV \cdot N program rollouts without minibatching, or M⋅N+V⋅(⌊N/5⌋+1)M \cdot N + V \cdot (\lfloor N/5 \rfloor + 1) with it: for N=10N = 10 and V=100V = 100 that is 350+300=650350 + 300 = 650 rollouts, plus the bootstrapping rollouts in step 1. Each rollout costs mm task-model calls. Dollars go almost entirely to the prompt model if the task model runs locally: about (prompt calls)⋅(tinpin+toutpout)(\text{prompt calls}) \cdot (t_{\text{in}} p_{\text{in}} + t_{\text{out}} p_{\text{out}}) with tt tokens per call and pp the price per token. Wall-clock time is dominated by the task-model rollouts, divided by num_threads.

Going further

  • GEPA: reflective prompt evolution, the optimizer that replaces the surrogate search with reflection on feedback.
  • Search and optimization basics, for exploring versus exploiting, which is what the Bayesian sampler balances.
  • The MIPRO paper by Opsahl-Ong and colleagues (2024), "Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs".
  • The source files dspy/teleprompt/mipro_optimizer_v2.py and dspy/propose/grounded_proposer.py.
  • Optuna's documentation on the tree-structured Parzen estimator, if you want the math behind step 3.

Back to DSPy and GEPA: programming and optimizing language-model systems