technique

Prompts and weights together

Combining prompt optimization and fine-tuning: alternating them with BetterTogether, ordering choices, and deciding which lever to pull first.

Before this

This page assumes you are comfortable with:

Why you need this

Stage 3 tunes the words a model reads; stage 4 tunes the model. They are usually presented as rivals, but they feed each other, and the order you run them in changes the result. This page shows how DSPy chains them with one optimizer, dspy.BetterTogether, and how to decide which lever to pull first for a given project.

The idea

Two quick definitions. Prompt optimization (few-shot bootstrapping, MIPROv2, GEPA) changes a program's instructions and worked examples, leaving the model's weights alone. Weight optimization (BootstrapFinetune, GRPO) changes the weights, the numbers inside the model.

Why they help each other:

  • Better prompts make better training data. BootstrapFinetune trains on the traces (the recorded inputs and outputs of every predictor) that pass your metric. A prompt-optimized program passes more often and passes in a more consistent shape, so there is more and cleaner data to train on. For GRPO, a better prompt means more groups with a mix of good and bad answers, which is where its learning signal comes from.
  • A tuned model reads prompts differently. After training, the instruction that suited the old model may be redundant or even misleading. A second round of prompt optimization adapts the instructions to the model you now have.

That gives the classic sequence: optimize prompts, then weights, then prompts again.

The BetterTogether paper by Soylu, Potts, and Khattab (2024) introduced this idea and reports improvements of up to 60% over optimizing weights alone and up to 6% over optimizing prompts alone, across three open 7-to-8-billion-parameter models and three tasks (multi-hop question answering, math reasoning, and classification). The dspy.BetterTogether source also cites a Databricks case study in which BetterTogether with GEPA and fine-tuning beat either approach alone; it gives no numbers in the source, so none are given here.

BetterTogether in DSPy 3.4

From the source of dspy.BetterTogether, exported at top level in DSPy 3.4:

Piece What it does
BetterTogether(metric, **optimizers) Each keyword becomes a name you can use in a strategy: p=dspy.GEPA(...), w=dspy.BootstrapFinetune(...). Any DSPy optimizer works.
Defaults With no optimizers given: p is BootstrapFewShotWithRandomSearch and w is BootstrapFinetune.
strategy A string of names joined by " -> " (space, arrow, space). Default "p -> w -> p".
valset Scores the program after every step. If you pass none, 10% of the trainset is held out (valset_ratio=0.1).
Result The best-scoring program on the valset (the latest one if there is no valset), with candidate_programs listing every step's program, score, and strategy so far, best first.
Failure If any step raises an error, it stops, sets flag_compilation_error_occurred, and returns the best program so far.
Models Every predictor must have its model set with set_lm, because weight optimizers need it.

Worked example

The strategy string, checked

The strategy is parsed before anything runs. This file builds a BetterTogether with GEPA as p and BootstrapFinetune as w, then calls the parsing step that compile runs first (a private method, used here only to show its behavior). GEPA's reflection model is DSPy's fake model, DummyLM, so nothing is called. Verified on DSPy 3.4.0 (ran on Python 3.14; 3.12 and newer behave the same).

import dspy
from dspy.utils import DummyLM


def metric(gold, pred, trace=None, pred_name=None, pred_trace=None):
    return float(gold.label == pred.label)


optimizer = dspy.BetterTogether(
    metric=metric,
    p=dspy.GEPA(metric=metric, auto="light", reflection_lm=DummyLM([])),
    w=dspy.BootstrapFinetune(metric=metric),
)
print(optimizer._prepare_strategy("p -> w -> p"))
try:
    optimizer._prepare_strategy("p -> w -> grpo")
except ValueError as err:
    print("ValueError:", err)
try:
    optimizer._prepare_strategy("p->w")
except ValueError as err:
    print("ValueError:", err)

Run with python strategy.py:

['p', 'w', 'p']
ValueError: Strategy contains invalid optimizer keys: ['grpo']. Valid keys are: ['p', 'w']
ValueError: Strategy contains invalid optimizer keys: ['p->w']. Valid keys are: ['p', 'w']

Two lessons. The spaces around the arrow are part of the separator. And the metric takes five arguments with defaults: dspy.GEPA refuses a metric that cannot accept (gold, pred, trace, pred_name, pred_trace), while BetterTogether's own scoring calls it with the first three. The defaults let one function serve both.

The full run

This sample needs a real model and training hardware; nothing shown was run. The student uses a model whose provider can fine-tune (see BootstrapFinetune for which ones can); GEPA's reflection model is Claude Opus 5.5 with its key read from the ANTHROPIC_API_KEY environment variable.

reflection = dspy.LM("anthropic/claude-opus-5-5")

optimizer = dspy.BetterTogether(
    metric=metric,
    p=dspy.GEPA(metric=metric, auto="light", reflection_lm=reflection),
    w=dspy.BootstrapFinetune(metric=metric, num_threads=16),
)

student.set_lm(student_lm)   # every predictor needs its model set
best = optimizer.compile(student, trainset=trainset, valset=valset, strategy="p -> w -> p")

for c in best.candidate_programs:
    print(c["strategy"] or "baseline", c["score"])

The loop prints one line per step plus the baseline, best first, so you can see whether the final p actually helped.

Which lever first

Your situation Pull first Why
Tens of labeled examples, not hundreds Prompts Too little data to train weights without memorizing it.
Few calls a day Prompts Training never pays for itself; see the cost formula on the BootstrapFinetune page.
No GPU and no hosted training Prompts Weight optimization is not available to you.
Many calls a day, a stable task, hundreds of examples or more Prompts, then weights ("p -> w", then try "p -> w -> p") Prompts first improve the training data; weights cut the cost per call.
Strict latency need Weights, eventually A small tuned model answers faster and needs fewer demos in its prompt.
Only a score, no teacher much better than the student, a multi-GPU training server Prompts, then GRPO GRPO learns from the score directly, but costs far more rollouts; run it last.

Three hypothetical projects

These are illustrations, not real projects.

A trivia bot for a club website. 60 labeled questions, a few hundred calls a week, one laptop. Every row of the table says prompts. Decision: run GEPA, ship the prompt-optimized program, stop. BetterTogether is not needed.

A support-ticket router for a busy shop. 3,000 labeled tickets, tens of thousands of calls a day, a GPU with enough memory to train a 1-billion-parameter model, and a need to answer in well under a second. Decision: BetterTogether with p=GEPA, w=BootstrapFinetune, strategy "p -> w -> p". Compare the candidates on the testset, not just the valset, and keep the strong-model prompt version as a fallback.

A math tutor graded by a rubric. 400 problems, a judge metric that gives partial credit, a moderate call volume, and access to a shared multi-GPU training server. Teacher traces are rarely fully right, so BootstrapFinetune's pass or fail filter keeps little. Decision: GEPA first; then, only if the gap to the target is still large, an experimental GRPO run as the w step. Budget for the GRPO step separately; it is where nearly all the cost goes.

In an optimization pipeline

This page is where stage 3 and stage 4 meet. Stage 2's metric and splits are shared by every step of a strategy, so a weak metric is amplified at every step. The valset picks the best step, which means the valset score of the winner is optimistic; the testset gives the honest number, as Overfitting and re-optimizing explains. Stage 5 must save both the prompts and the tuned model the program now points to.

Common mistakes

  • Training weights before prompts. BootstrapFinetune then trains on the unoptimized teacher's traces, which pass the metric less often, so there are fewer of them. If you want to know, run both orders and compare their candidate_programs scores.
  • Forgetting set_lm. The student must carry its own model on every predictor. Without it, compile fails at the baseline evaluation with AttributeError: 'NoneType' object has no attribute 'launch', an error that does not mention set_lm.
  • Writing the strategy as "p->w". It is read as one unknown optimizer name, and compile stops before doing anything.
  • Ignoring the error flag. A failed step does not raise; it returns the best program so far. Check best.flag_compilation_error_occurred after every run.
  • Reporting the valset score. The winning step was chosen on the valset, so its score is biased upward. Report the testset.
  • A 10% held-out valset from a small trainset. With 100 examples, that is 10 validation examples, and the "best" step is decided by noise. Pass a real valset.

Cost

A strategy costs the sum of its steps plus one valset evaluation before the first step and after each one. With mm steps, valset size vv, and kk predictor calls per rollout, the evaluations alone are (m+1)vk(m + 1) v k calls to the student model. Each prompt step costs what GEPA or MIPROv2 costs on its own (GEPA in practice gives the formula). Each weight step costs a BootstrapFinetune data collection (nkn k teacher calls for a trainset of nn) plus training hardware, or a GRPO run. For "p -> w -> p" with v=100v = 100 and k=2k = 2, the evaluations are 4×100×2=8004 \times 100 \times 2 = 800 student calls, small next to the GEPA and training steps. The builder's time goes into the training stack and into reading the candidate list after each run.

Going further

  • The BetterTogether paper by Soylu, Potts, and Khattab, "Fine-Tuning and Prompt Optimization: Two Great Steps that Work Better Together".
  • The dspy.BetterTogether source, whose docstring shows GEPA, MIPROv2, and BootstrapFinetune combinations.
  • The GEPA paper by Agrawal and colleagues, for why prompt optimization is so much cheaper per point gained.
  • Saving and serving programs, for keeping the result.

Back to DSPy and GEPA: programming and optimizing language-model systems