technique

Composing programs

Building a multi-step program as a dspy.Module subclass: several predictors wired together in forward, with tools, loops, and branches in plain Python.

Before this

This page assumes you are comfortable with:

Why you need this

Most real tasks are more than one model call. You pull the relevant facts out of a document, then answer from them; you look something up, then decide. Writing all of that as one giant prompt makes it hard to see which step failed and impossible to improve one step without disturbing the others. DSPy lets you write each step as its own predictor and wire them together in ordinary Python, and every optimizer in this cluster then works on each step separately.

The idea

A program in DSPy is a Python class that inherits from dspy.Module. It has two methods you write:

  • __init__ declares the parts: each sub-module is stored as an attribute, such as self.extract = dspy.Predict(...).
  • forward says how the parts connect. It takes the program's inputs as keyword arguments, calls the parts, and returns a dspy.Prediction, which is an object holding named output fields.

You never call forward directly. You call the program like a function, program(question=...), and dspy.Module runs forward for you, with callbacks and usage tracking around it. Inside forward, anything Python can do is allowed: if statements, loops, calls to your own functions, a database lookup. The model calls are only the predictors.

Predictors and their names

A predictor is one dspy.Predict inside the program: one signature, one instruction, its demos (worked examples), one model call each time it runs. Modules like dspy.ChainOfThought are wrappers that contain a predictor. program.named_predictors() walks the program and lists every predictor with a name built from the attribute path: extract for self.extract, and respond.predict for the predictor inside self.respond = dspy.ChainOfThought(...).

These names are not cosmetic. In the DSPy 3.4.0 source:

Who uses the name How
dspy.GEPA A candidate is a dictionary from predictor name to instruction text, built from named_predictors(). The default component_selector="round_robin" picks which name to rewrite next.
A GEPA feedback metric Receives pred_name, the predictor being improved, so it can give step-specific feedback.
program.save(...) The saved state is keyed by these names. Rename an attribute and an old save no longer loads into it.

So name attributes for what the step does (extract, respond), and treat a rename after optimization as a breaking change.

Tools with ReAct

dspy.ReAct(signature, tools, max_iters=20) is a built-in module for a loop where the model picks a tool, sees the result, and repeats until it chooses finish. A tool is a plain Python function; its name, type hints, and docstring become the tool description the model reads. In the installed 3.4.0 source, ReAct holds two predictors: react, which writes next_thought, next_tool_name, and next_tool_args on each turn, and extract, a ChainOfThought that reads the whole trajectory and writes the final outputs. The DSPy ReAct documentation also describes an experimental dspy.ReActV2; this page sticks to dspy.ReAct, the current one through 3.4.

Worked example

A two-step program: pull the facts that matter out of a passage, then answer from those facts only. It also branches: if no facts are found, it answers without a second model call. All samples were run on DSPy 3.4.0 (on Python 3.14; 3.12 and newer behave the same) against DummyLM, the fake LM that returns scripted answers; in real use the local task model (a model served by Ollama) takes its place.

# two_step.py
import dspy
from dspy.utils import DummyLM


class ExtractFacts(dspy.Signature):
    """List the facts in the passage that matter for the question."""

    passage: str = dspy.InputField()
    question: str = dspy.InputField()
    facts: list[str] = dspy.OutputField()


class AnswerFromFacts(dspy.Signature):
    """Answer the question using only the listed facts."""

    question: str = dspy.InputField()
    facts: list[str] = dspy.InputField()
    answer: str = dspy.OutputField()


class FactsThenAnswer(dspy.Module):
    def __init__(self):
        super().__init__()
        self.extract = dspy.Predict(ExtractFacts)
        self.respond = dspy.ChainOfThought(AnswerFromFacts)

    def forward(self, passage, question):
        found = self.extract(passage=passage, question=question)
        if not found.facts:
            return dspy.Prediction(facts=[], answer="Not stated in the passage.")
        result = self.respond(question=question, facts=found.facts)
        return dspy.Prediction(facts=found.facts, answer=result.answer)


lm = DummyLM([
    {"facts": ["The bridge opened in 1932.", "It is 503 meters long."]},
    {"reasoning": "The second fact gives the length.", "answer": "503 meters"},
])
dspy.configure(lm=lm)

program = FactsThenAnswer()
for name, predictor in program.named_predictors():
    print(name, "->", type(predictor).__name__)

pred = program(
    passage="The bridge opened in 1932. It is 503 meters long and painted grey.",
    question="How long is the bridge?",
)
print(pred.facts)
print(pred.answer)
print("model calls:", len(lm.history))

Run with tmp/dspy-venv/Scripts/python.exe two_step.py:

extract -> Predict
respond.predict -> Predict
['The bridge opened in 1932.', 'It is 503 meters long.']
503 meters
model calls: 2

Step by step:

Step Code Model call? Result
1 self.extract(...) Yes, predictor extract Two facts, parsed into a Python list[str]
2 if not found.facts No, plain Python Facts exist, so continue
3 self.respond(...) Yes, predictor respond.predict Reasoning, then answer
4 return dspy.Prediction(...) No facts and answer together

The ChainOfThought added a reasoning output field ahead of answer, which is why the second scripted answer has one. The program returned only the fields it chose to.

ReAct with one tool

# react_tool.py
import dspy
from dspy.utils import DummyLM

STOCK = {"red mug": 4, "blue mug": 0}


def check_stock(item: str) -> int:
    """Return how many of an item are in the warehouse."""
    return STOCK.get(item, 0)


agent = dspy.ReAct("question -> answer", tools=[check_stock], max_iters=5)

lm = DummyLM([
    {"next_thought": "I should look up the blue mug.",
     "next_tool_name": "check_stock",
     "next_tool_args": {"item": "blue mug"}},
    {"next_thought": "The count is 0, so I can answer.",
     "next_tool_name": "finish",
     "next_tool_args": {}},
    {"reasoning": "The tool returned 0.", "answer": "No, the blue mug is out of stock."},
])
dspy.configure(lm=lm)

for name, predictor in agent.named_predictors():
    print(name, "->", type(predictor).__name__)

pred = agent(question="Is the blue mug in stock?")
for key, value in pred.trajectory.items():
    print(f"{key}: {value}")
print("answer:", pred.answer)
print("model calls:", len(lm.history))

Output:

react -> Predict
extract.predict -> Predict
thought_0: I should look up the blue mug.
tool_name_0: check_stock
tool_args_0: {'item': 'blue mug'}
observation_0: 0
thought_1: The count is 0, so I can answer.
tool_name_1: finish
tool_args_1: {}
observation_1: Completed.
answer: No, the blue mug is out of stock.
model calls: 3

The tool ran for real: observation_0: 0 came from check_stock, not from the script. The scripted parts are the model's choices. Three model calls answered one question: two turns of react, one of extract.

In an optimization pipeline

This is stage 1, defining the program. Its shape decides what stages 3 and 4 can do. dspy.BootstrapFewShot collects demos for every predictor, including extract, whose outputs nobody labeled, by keeping traces (the record of each predictor's inputs and outputs during a rollout) from runs that passed the metric. dspy.GEPA rewrites one predictor's instruction at a time, by name. With use_merge=True (the default) it also tries merging two good candidates: the merge step in the installed gepa 0.1.4 source takes each predictor's instruction from whichever parent changed it since their common ancestor. That only makes sense when the program has several predictors. Fine-tuning in stage 4 builds training data per predictor.

Common mistakes

  • Returning a plain dict from forward. Metrics and dspy.Evaluate read fields as attributes (pred.answer), so a dict fails with AttributeError: 'dict' object has no attribute 'answer' the first time it is scored. Return dspy.Prediction(...).
  • Creating predictors inside forward. A dspy.Predict made inside forward is new on every call, so optimizers never see it and its improved instruction is thrown away. Symptom: an optimized program behaves exactly like the original.
  • Hiding predictors where the walk cannot see them. named_predictors() looks at attributes, and one level into lists, tuples, and dicts. A predictor in a list inside a list is skipped. Symptom: GEPA never rewrites that step, and it is missing from the saved file.
  • Calling program.forward(...) directly. It works, but skips what __call__ adds around it, such as callbacks and usage tracking. Call program(...).
  • Unbounded tool loops. A ReAct agent that never picks finish runs max_iters turns (20 by default). Symptom: one question costs 21 model calls. Set max_iters to what the task needs.

Cost

A program with kk predictors makes about kk model calls per rollout (a rollout is one run on one example) when there are no loops. ReAct makes up to m+1m + 1, where mm is max_iters: one per turn plus the final extraction. The ReAct trajectory is resent on every turn, so input tokens grow with each turn: if each turn adds about tt tokens, turn jj sends roughly t0+jtt_0 + jt tokens, where t0t_0 is the starting prompt, and the whole run sends about mt0+t m(m−1)/2m t_0 + t\,m(m-1)/2 input tokens. Optimizer cost multiplies by these call counts, so a three-step program costs about three times as much to optimize as a one-step one on the same budget of rollouts. Builder time is the gain: a failing step shows up by name in inspect_history and in GEPA's per-predictor feedback.

Going further

Leads to

Back to DSPy and GEPA: programming and optimizing language-model systems