technique

Signatures and modules

Writing a task as typed inputs and outputs instead of a prompt string, and running it through Predict, ChainOfThought, and other built-in modules.

Before this

This page assumes you are comfortable with:

Why you need this

A hand-written prompt mixes three things in one string: what the task is, how to phrase it for one particular model, and which examples to show. When the model changes, all three get edited together, by hand, with no measurement. DSPy, a Python framework for building systems around language models, separates them. You write the task as a signature (typed inputs and outputs) and pick a module (a strategy for calling the model). DSPy builds the actual prompt from those, and later an optimizer tunes the parts of the prompt that are meant to change. Every sample below ran on DSPy 3.4 (3.4.0) against DummyLM, DSPy's fake language model, which returns scripted answers instead of calling a real model.

The idea

Program, not prompt

The core idea of DSPy: the thing you write and keep is a program, a tree of modules. The prompt is an output, generated from the program each time it runs, the way a compiler generates machine code from source. Change the model, and DSPy generates a prompt for the new one from the same program. Run an optimizer, and it changes the program's adjustable parts, then the prompt follows.

Signatures

A signature declares a task's input fields and output fields by name and type. The simplest form is an inline signature, a string with an arrow:

"review -> sentiment"

That means one input field named review and one output field named sentiment, both strings. Add types with a colon and separate fields with commas: "review -> positive: bool, stars: int". Field names matter: DSPy puts them in the prompt, so review and sentiment tell the model a great deal by themselves.

An inline signature gets a default instruction made from its field names, such as Given the fields `review`, produce the fields `sentiment`. To write your own, pass instructions= to dspy.Signature.

A class-based signature is a Python class that subclasses dspy.Signature. Its docstring becomes the instruction, and each field is a type-hinted attribute set to dspy.InputField() or dspy.OutputField(), which accept an optional desc= with a short description:

# excerpt of sentiment.py, shown in full under Worked example
class Sentiment(dspy.Signature):
    """Classify the sentiment of a product review."""

    review: str = dspy.InputField()
    sentiment: Literal["positive", "negative", "neutral"] = dspy.OutputField()

The type hint on an output field is a promise DSPy enforces. Literal[...] limits the answer to listed strings; bool, int, float, list[str], and Pydantic models also work. DSPy tells the model the expected type in the prompt and converts the reply into that Python type, or raises an error if it cannot.

Modules

A module takes a signature and decides how to call the model with it. You create one, then call it with keyword arguments named after the input fields. It returns a dspy.Prediction, an object with one attribute per output field.

  • dspy.Predict(signature) is the basic module: one model call, signature unchanged.
  • dspy.ChainOfThought(signature) adds an output field named reasoning before your outputs, so the model writes its reasoning first. Inside, it is a dspy.Predict on the extended signature, stored as the attribute predict.

Other built-in modules in DSPy 3.4, checked against the installed source:

Module What it does
dspy.ReAct A tool-using agent: reasons, calls Python functions you supply, and repeats until it can fill the outputs. dspy.ReActV2 is a newer variant with native function calling.
dspy.RLM A "recursive language model" that explores a large input by writing code in a sandboxed Python interpreter.
dspy.ProgramOfThought Has the model write Python whose result is the answer. Deprecated in 3.4 and scheduled for removal in 3.5; the source names dspy.RLM as the replacement.
dspy.MultiChainComparison Compares several reasoning attempts and produces one final answer.
dspy.BestOfN Runs a module up to N times with a reward function you write and returns the best attempt.
dspy.Refine Like BestOfN, but after a weak attempt it generates feedback to improve the next one.
dspy.Parallel Runs a module over many inputs on several threads.

What a module can learn

Each dspy.Predict inside a program is a predictor: one model call. A predictor carries two adjustable parts, the program's learnable parameters:

Parameter Where it lives Starts as
Instruction predictor.signature.instructions Your docstring, or the default sentence
Demos (worked examples) predictor.demos An empty list

Prompt optimizers change exactly these two things; the type hints and field names stay as you wrote them. (Weight optimizers, in stage 4, change the model instead.) These are the instruction and the examples from Prompts and few-shot learning, now held as data that a program can edit.

Worked example

A sentiment signature run through Predict and ChainOfThought. The fake model is scripted with one answer per call, in order.

# sentiment.py
import sys
from typing import Literal

import dspy
from dspy.utils import DummyLM


class Sentiment(dspy.Signature):
    """Classify the sentiment of a product review."""

    review: str = dspy.InputField()
    sentiment: Literal["positive", "negative", "neutral"] = dspy.OutputField()


lm = DummyLM([
    {"sentiment": "positive"},
    {"reasoning": "The reviewer praises the battery life.", "sentiment": "positive"},
])
dspy.configure(lm=lm)

review = "Battery lasts two days. Love it."

predict = dspy.Predict(Sentiment)
cot = dspy.ChainOfThought(Sentiment)

p = predict(review=review)
print("Predict:", p.sentiment)

c = cot(review=review)
print("ChainOfThought:", c.sentiment)
print("Reasoning:", c.reasoning)

for name, predictor in cot.named_predictors():
    print("Predictor:", name)
    print("  fields:", list(predictor.signature.fields))
    print("  instructions:", predictor.signature.instructions)
    print("  demos:", predictor.demos)

dspy.inspect_history(n=1, file=sys.stdout)

dspy.configure(lm=...) sets the model every module uses (Configuring language models). dspy.inspect_history(n=1) prints the last model call; passing file=sys.stdout turns off its terminal colors. python sentiment.py prints:

Predict: positive
ChainOfThought: positive
Reasoning: The reviewer praises the battery life.
Predictor: predict
  fields: ['review', 'reasoning', 'sentiment']
  instructions: Classify the sentiment of a product review.
  demos: []




[2026-10-03T09:12:06.718648]

System message:

Your input fields are:
1. `review` (str):
Your output fields are:
1. `reasoning` (str): 
2. `sentiment` (Literal['positive', 'negative', 'neutral']):
All interactions will be structured in the following way, with the appropriate values filled in.

[[ ## review ## ]]
{review}

[[ ## reasoning ## ]]
{reasoning}

[[ ## sentiment ## ]]
{sentiment}        # note: the value you produce must exactly match (no extra characters) one of: positive; negative; neutral

[[ ## completed ## ]]
In adhering to this structure, your objective is: 
        Classify the sentiment of a product review.


User message:

[[ ## review ## ]]
Battery lasts two days. Love it.

Respond with the corresponding output fields, starting with the field `[[ ## reasoning ## ]]`, then `[[ ## sentiment ## ]]` (must be formatted as a valid Python Literal['positive', 'negative', 'neutral']), and then ending with the marker for `[[ ## completed ## ]]`.


Response:

[[ ## reasoning ## ]]
The reviewer praises the battery life.

[[ ## sentiment ## ]]
positive

The answers are staged, but the prompt is the real one DSPy 3.4 generates. Read it against the program:

In the prompt Came from
The two lists of fields and their types The signature's fields, with reasoning added by ChainOfThought
[[ ## review ## ]] style markers DSPy's default formatter (the ChatAdapter), so the reply can be split back into fields
must exactly match ... one of: positive; negative; neutral The Literal type hint
your objective is: Classify the sentiment... The docstring, which is the instruction
No example conversations demos is empty; an optimizer would fill it

There is one predictor, named predict after the attribute inside ChainOfThought. Its instruction and its empty demos list are the two things an optimizer would change.

The inline form works the same way, and typed fields come back as Python types:

# inline.py
import dspy
from dspy.utils import DummyLM

lm = DummyLM([{"sentiment": "negative"}, {"positive": "true", "stars": "4"}])
dspy.configure(lm=lm)

classify = dspy.Predict("review -> sentiment")
result = classify(review="Arrived broken.")
print(result)
print(classify.signature.instructions)

rate = dspy.Predict(dspy.Signature("review -> positive: bool, stars: int",
                                   instructions="Rate the review from 1 to 5 stars."))
r = rate(review="Battery lasts two days. Love it.")
print(r.positive, type(r.positive).__name__, r.stars, type(r.stars).__name__)
print(rate.signature.instructions)

python inline.py prints:

Prediction(
    sentiment='negative'
)
Given the fields `review`, produce the fields `sentiment`.
True bool 4 int
Rate the review from 1 to 5 stars.

The model replied with the text true and 4; the program received the Python values True and 4.

In an optimization pipeline

This is stage 1, Define the program. Signatures and modules are what every later stage works on: Composing programs wires several predictors together in one module, Adapters and structured output explains the field markers and the parsing, stage 2 runs the program over a dataset and scores it, and stage 3's optimizers edit each predictor's instruction and demos. Because the program is code and the prompt is generated, an optimized program can be saved as data and re-run against a different model.

Common mistakes

  • Calling a module with positional arguments. predict("Arrived broken.") raises a ValueError saying positional arguments are not allowed and naming the expected keyword, review.
  • Input names that do not match the signature. Calling predict(text=...) only logs two warnings (text will be ignored, review is missing). The call still runs: in the fake-model run, the prompt contained no review at all, and an answer came back anyway. A real model would guess.
  • Vague field names. x -> y gives the model nothing to go on; review -> sentiment already says most of the task.
  • No type on a label field. Without Literal[...] the model can answer Positive! or mostly positive, and an exact-match metric scores it wrong.
  • A reply that does not fit the type. When the fake model answered great for the Literal field, parsing failed, DSPy retried once with a JSON format (its fallback), and the call ended in AdapterParseError. With a real model, each retry is another paid call.
  • Reading result.reasoning from dspy.Predict. Only ChainOfThought adds that field; on a Predict result it raises AttributeError: 'Prediction' object has no attribute 'reasoning'.

Cost

A dspy.Predict call is one model call. Its prompt is the field descriptions and markers, the instruction, every demo, and the input; with kk demos of tEt_E tokens each, the demos alone add k⋅tEk \cdot t_E input tokens to every call. dspy.ChainOfThought is also one call, but it generates the reasoning tokens before the answer, adding output tokens (often the more expensive kind) and time. BestOfN and Refine make up to NN calls per input. A parse failure that triggers the JSON fallback makes a second call. Writing a signature costs minutes; the payoff is that the prompt no longer needs rewriting by hand when the model changes.

Going further

  • The DSPy Signatures and Modules documentation pages.
  • Composing programs, for several predictors in one dspy.Module.
  • Adapters and structured output, for ChatAdapter, JSONAdapter, and the fallback.
  • The DSPy paper by Khattab and colleagues (2023), "DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines".

Leads to

Back to DSPy and GEPA: programming and optimizing language-model systems