technique
Composing programs
Building a multi-step program as a dspy.Module subclass: several predictors wired together in forward, with tools, loops, and branches in plain Python.
Before this
This page assumes you are comfortable with:
Why you need this
Most real tasks are more than one model call. You pull the relevant facts out of a document, then answer from them; you look something up, then decide. Writing all of that as one giant prompt makes it hard to see which step failed and impossible to improve one step without disturbing the others. DSPy lets you write each step as its own predictor and wire them together in ordinary Python, and every optimizer in this cluster then works on each step separately.
The idea
A program in DSPy is a Python class that inherits from dspy.Module. It has two methods you write:
__init__declares the parts: each sub-module is stored as an attribute, such asself.extract = dspy.Predict(...).forwardsays how the parts connect. It takes the program's inputs as keyword arguments, calls the parts, and returns adspy.Prediction, which is an object holding named output fields.
You never call forward directly. You call the program like a function, program(question=...), and dspy.Module runs forward for you, with callbacks and usage tracking around it. Inside forward, anything Python can do is allowed: if statements, loops, calls to your own functions, a database lookup. The model calls are only the predictors.
Predictors and their names
A predictor is one dspy.Predict inside the program: one signature, one instruction, its demos (worked examples), one model call each time it runs. Modules like dspy.ChainOfThought are wrappers that contain a predictor. program.named_predictors() walks the program and lists every predictor with a name built from the attribute path: extract for self.extract, and respond.predict for the predictor inside self.respond = dspy.ChainOfThought(...).
These names are not cosmetic. In the DSPy 3.4.0 source:
| Who uses the name | How |
|---|---|
dspy.GEPA |
A candidate is a dictionary from predictor name to instruction text, built from named_predictors(). The default component_selector="round_robin" picks which name to rewrite next. |
| A GEPA feedback metric | Receives pred_name, the predictor being improved, so it can give step-specific feedback. |
program.save(...) |
The saved state is keyed by these names. Rename an attribute and an old save no longer loads into it. |
So name attributes for what the step does (extract, respond), and treat a rename after optimization as a breaking change.
Tools with ReAct
dspy.ReAct(signature, tools, max_iters=20) is a built-in module for a loop where the model picks a tool, sees the result, and repeats until it chooses finish. A tool is a plain Python function; its name, type hints, and docstring become the tool description the model reads. In the installed 3.4.0 source, ReAct holds two predictors: react, which writes next_thought, next_tool_name, and next_tool_args on each turn, and extract, a ChainOfThought that reads the whole trajectory and writes the final outputs. The DSPy ReAct documentation also describes an experimental dspy.ReActV2; this page sticks to dspy.ReAct, the current one through 3.4.
Worked example
A two-step program: pull the facts that matter out of a passage, then answer from those facts only. It also branches: if no facts are found, it answers without a second model call. All samples were run on DSPy 3.4.0 (on Python 3.14; 3.12 and newer behave the same) against DummyLM, the fake LM that returns scripted answers; in real use the local task model (a model served by Ollama) takes its place.
# two_step.py
import dspy
from dspy.utils import DummyLM
class ExtractFacts(dspy.Signature):
"""List the facts in the passage that matter for the question."""
passage: str = dspy.InputField()
question: str = dspy.InputField()
facts: list[str] = dspy.OutputField()
class AnswerFromFacts(dspy.Signature):
"""Answer the question using only the listed facts."""
question: str = dspy.InputField()
facts: list[str] = dspy.InputField()
answer: str = dspy.OutputField()
class FactsThenAnswer(dspy.Module):
def __init__(self):
super().__init__()
self.extract = dspy.Predict(ExtractFacts)
self.respond = dspy.ChainOfThought(AnswerFromFacts)
def forward(self, passage, question):
found = self.extract(passage=passage, question=question)
if not found.facts:
return dspy.Prediction(facts=[], answer="Not stated in the passage.")
result = self.respond(question=question, facts=found.facts)
return dspy.Prediction(facts=found.facts, answer=result.answer)
lm = DummyLM([
{"facts": ["The bridge opened in 1932.", "It is 503 meters long."]},
{"reasoning": "The second fact gives the length.", "answer": "503 meters"},
])
dspy.configure(lm=lm)
program = FactsThenAnswer()
for name, predictor in program.named_predictors():
print(name, "->", type(predictor).__name__)
pred = program(
passage="The bridge opened in 1932. It is 503 meters long and painted grey.",
question="How long is the bridge?",
)
print(pred.facts)
print(pred.answer)
print("model calls:", len(lm.history))
Run with tmp/dspy-venv/Scripts/python.exe two_step.py:
extract -> Predict
respond.predict -> Predict
['The bridge opened in 1932.', 'It is 503 meters long.']
503 meters
model calls: 2
Step by step:
| Step | Code | Model call? | Result |
|---|---|---|---|
| 1 | self.extract(...) |
Yes, predictor extract |
Two facts, parsed into a Python list[str] |
| 2 | if not found.facts |
No, plain Python | Facts exist, so continue |
| 3 | self.respond(...) |
Yes, predictor respond.predict |
Reasoning, then answer |
| 4 | return dspy.Prediction(...) |
No | facts and answer together |
The ChainOfThought added a reasoning output field ahead of answer, which is why the second scripted answer has one. The program returned only the fields it chose to.
ReAct with one tool
# react_tool.py
import dspy
from dspy.utils import DummyLM
STOCK = {"red mug": 4, "blue mug": 0}
def check_stock(item: str) -> int:
"""Return how many of an item are in the warehouse."""
return STOCK.get(item, 0)
agent = dspy.ReAct("question -> answer", tools=[check_stock], max_iters=5)
lm = DummyLM([
{"next_thought": "I should look up the blue mug.",
"next_tool_name": "check_stock",
"next_tool_args": {"item": "blue mug"}},
{"next_thought": "The count is 0, so I can answer.",
"next_tool_name": "finish",
"next_tool_args": {}},
{"reasoning": "The tool returned 0.", "answer": "No, the blue mug is out of stock."},
])
dspy.configure(lm=lm)
for name, predictor in agent.named_predictors():
print(name, "->", type(predictor).__name__)
pred = agent(question="Is the blue mug in stock?")
for key, value in pred.trajectory.items():
print(f"{key}: {value}")
print("answer:", pred.answer)
print("model calls:", len(lm.history))
Output:
react -> Predict
extract.predict -> Predict
thought_0: I should look up the blue mug.
tool_name_0: check_stock
tool_args_0: {'item': 'blue mug'}
observation_0: 0
thought_1: The count is 0, so I can answer.
tool_name_1: finish
tool_args_1: {}
observation_1: Completed.
answer: No, the blue mug is out of stock.
model calls: 3
The tool ran for real: observation_0: 0 came from check_stock, not from the script. The scripted parts are the model's choices. Three model calls answered one question: two turns of react, one of extract.
In an optimization pipeline
This is stage 1, defining the program. Its shape decides what stages 3 and 4 can do. dspy.BootstrapFewShot collects demos for every predictor, including extract, whose outputs nobody labeled, by keeping traces (the record of each predictor's inputs and outputs during a rollout) from runs that passed the metric. dspy.GEPA rewrites one predictor's instruction at a time, by name. With use_merge=True (the default) it also tries merging two good candidates: the merge step in the installed gepa 0.1.4 source takes each predictor's instruction from whichever parent changed it since their common ancestor. That only makes sense when the program has several predictors. Fine-tuning in stage 4 builds training data per predictor.
Common mistakes
- Returning a plain dict from
forward. Metrics anddspy.Evaluateread fields as attributes (pred.answer), so a dict fails withAttributeError: 'dict' object has no attribute 'answer'the first time it is scored. Returndspy.Prediction(...). - Creating predictors inside
forward. Adspy.Predictmade insideforwardis new on every call, so optimizers never see it and its improved instruction is thrown away. Symptom: an optimized program behaves exactly like the original. - Hiding predictors where the walk cannot see them.
named_predictors()looks at attributes, and one level into lists, tuples, and dicts. A predictor in a list inside a list is skipped. Symptom: GEPA never rewrites that step, and it is missing from the saved file. - Calling
program.forward(...)directly. It works, but skips what__call__adds around it, such as callbacks and usage tracking. Callprogram(...). - Unbounded tool loops. A ReAct agent that never picks
finishrunsmax_itersturns (20 by default). Symptom: one question costs 21 model calls. Setmax_itersto what the task needs.
Cost
A program with predictors makes about model calls per rollout (a rollout is one run on one example) when there are no loops. ReAct makes up to , where is max_iters: one per turn plus the final extraction. The ReAct trajectory is resent on every turn, so input tokens grow with each turn: if each turn adds about tokens, turn sends roughly tokens, where is the starting prompt, and the whole run sends about input tokens. Optimizer cost multiplies by these call counts, so a three-step program costs about three times as much to optimize as a one-step one on the same budget of rollouts. Builder time is the gain: a failing step shows up by name in inspect_history and in GEPA's per-predictor feedback.
Going further
- Adapters and structured output: how each predictor's messages are built.
- Writing metrics: using
pred_nameto give each step its own feedback. - GEPA: reflective prompt evolution: round-robin predictor selection and system-aware merge.
- The DSPy ReAct documentation, including the experimental
dspy.ReActV2.
Leads to
- techniqueFew-shot bootstrappingThe first family of optimizers: LabeledFewShot, BootstrapFewShot, and random search over bootstrapped demos, which pick worked examples for every predictor automatically.
- techniqueGEPA: reflective prompt evolutionHow GEPA improves prompts by reading what went wrong: run a minibatch, reflect on traces and feedback in plain language, propose a new instruction, and keep a Pareto front of candidates.
Back to DSPy and GEPA: programming and optimizing language-model systems