technique

Adapters and structured output

How DSPy turns a signature into the actual messages a model sees and parses the reply back into typed fields: ChatAdapter, JSONAdapter, and what to do when parsing fails.

Before this

This page assumes you are comfortable with:

Why you need this

A signature says "a review goes in, a sentiment label and a star count come out". A language model, though, only reads and writes text. Something has to turn the signature into messages, and something has to turn the model's reply back into a Python string and a Python int. In DSPy that something is the adapter. When a program fails with a parse error, or an optimizer's rewritten instruction suddenly breaks the output format, the adapter is where you look.

The idea

An adapter is the layer between a signature and the language model (LM). It does two jobs on every call:

  1. Format. Build a list of chat messages: a system message that lists the fields, their types, and the instruction, then a user message holding the input values.
  2. Parse. Read the model's reply, find each output field, and convert it to the type the signature declared (str, int, bool, list[str], a Literal set of labels, or a Pydantic model, which is a Python class describing a JSON object's fields).

DSPy 3.4.0, the version every sample on this page was run on (on Python 3.14; 3.12 and newer behave the same), exports four adapters at the top level: dspy.ChatAdapter, dspy.JSONAdapter, dspy.XMLAdapter, and dspy.TwoStepAdapter. This page covers the first two, which are the ones you will meet in practice.

dspy.ChatAdapter dspy.JSONAdapter
Default? Yes, when no adapter is configured No, you choose it
Reply format it asks for Each field under a marker line [[ ## name ## ]], ending with [[ ## completed ## ]] One JSON object with a key per output field
Provider structured output Not requested Requested when the LM supports a response schema; otherwise plain JSON mode
On a parse failure Retries once with JSONAdapter (on by default) Raises the error

You pick an adapter for the whole program with dspy.configure(adapter=dspy.JSONAdapter()), or for a block of code with with dspy.context(adapter=...).

Typed outputs

The type on an output field does real work. It is written into the prompt as a note (for example "must be a single int value"), and the parser checks the reply against it. A Literal["positive", "negative", "neutral"] field is the cheapest way to get a label from a fixed set: anything else is a parse error, not a silent bad value. A Pydantic model as a field type puts its JSON schema straight into the prompt; in a check run for this page, an Order model with item: str and quantity: int produced this note:

{order}        # note: the value you produce must adhere to the JSON schema: {"type": "object", "properties": {"item": {"type": "string", "title": "Item"}, "quantity": {"type": "integer", "title": "Quantity"}}, "required": ["item", "quantity"], "title": "Order"}

and the reply {"item": "blue mug", "quantity": 3} came back as Order(item='blue mug', quantity=3), a real Python object.

The fallback rule

From the ChatAdapter source in 3.4.0: when parsing raises AdapterParseError, ChatAdapter makes one more call with a JSONAdapter, unless any of these is true: fallback was turned off with ChatAdapter(use_json_adapter_fallback=False), the adapter already is a JSONAdapter, or part of the reply was already streamed to the user. Only a parse failure triggers it. A network error, a missing API key, or a bug in your code goes straight to you.

Worked example

One signature, two adapters, the same scripted answer. DummyLM is DSPy's fake LM: it returns the answers you give it, formatted the way the chosen adapter expects, and makes no network call. In real use the local task model (a model served by Ollama) takes its place.

# two_adapters.py
from typing import Literal

import dspy
from dspy.utils import DummyLM


class ReviewLabel(dspy.Signature):
    """Label the sentiment of a product review."""

    review: str = dspy.InputField()
    sentiment: Literal["positive", "negative", "neutral"] = dspy.OutputField()
    stars: int = dspy.OutputField(desc="1 to 5")


classify = dspy.Predict(ReviewLabel)
answer = {"sentiment": "negative", "stars": 2}

for adapter in [dspy.ChatAdapter(), dspy.JSONAdapter()]:
    lm = DummyLM([answer], adapter=adapter)
    with dspy.context(lm=lm, adapter=adapter):
        pred = classify(review="Broke after two days.")
    print(type(adapter).__name__, "->", repr(pred.sentiment), repr(pred.stars))
    dspy.inspect_history(n=1)

Run with tmp/dspy-venv/Scripts/python.exe two_adapters.py. dspy.inspect_history(n=1) prints the last call's messages. Output, with terminal colors removed and blank lines trimmed:

ChatAdapter -> 'negative' 2

System message:

Your input fields are:
1. `review` (str):
Your output fields are:
1. `sentiment` (Literal['positive', 'negative', 'neutral']): 
2. `stars` (int): 1 to 5
All interactions will be structured in the following way, with the appropriate values filled in.

[[ ## review ## ]]
{review}

[[ ## sentiment ## ]]
{sentiment}        # note: the value you produce must exactly match (no extra characters) one of: positive; negative; neutral

[[ ## stars ## ]]
{stars}        # note: the value you produce must be a single int value

[[ ## completed ## ]]
In adhering to this structure, your objective is: 
        Label the sentiment of a product review.

User message:

[[ ## review ## ]]
Broke after two days.

Respond with the corresponding output fields, starting with the field `[[ ## sentiment ## ]]` (must be formatted as a valid Python Literal['positive', 'negative', 'neutral']), then `[[ ## stars ## ]]` (must be formatted as a valid Python int), and then ending with the marker for `[[ ## completed ## ]]`.

Response:

[[ ## sentiment ## ]]
negative

[[ ## stars ## ]]
2

JSONAdapter -> 'negative' 2

System message:
...
All interactions will be structured in the following way, with the appropriate values filled in.

Inputs will have the following structure:

[[ ## review ## ]]
{review}

Outputs will be a JSON object with the following fields.

{
  "sentiment": "{sentiment}        # note: the value you produce must exactly match (no extra characters) one of: positive; negative; neutral",
  "stars": "{stars}        # note: the value you produce must be a single int value"
}
In adhering to this structure, your objective is: 
        Label the sentiment of a product review.

User message:

[[ ## review ## ]]
Broke after two days.

Respond with a JSON object in the following order of fields: `sentiment` (must be formatted as a valid Python Literal['positive', 'negative', 'neutral']), then `stars` (must be formatted as a valid Python int).

Response:

{
  "sentiment": "negative",
  "stars": 2
}

What to notice, part by part:

Part ChatAdapter JSONAdapter
Field list Same in both Same in both
Inputs Marker lines Marker lines too (only outputs change)
Outputs requested as Marker lines, then [[ ## completed ## ]] One JSON object
Instruction The docstring, after "your objective is:" Same
Python result 'negative', 2 'negative', 2

The docstring "Label the sentiment of a product review." is the instruction. It is the text GEPA and MIPROv2 rewrite; everything else in the system message is generated from the field names and types. The scripted ChatAdapter reply has no [[ ## completed ## ]] line, and it still parsed: the parser looks for the output fields, not the end marker.

When parsing fails

A second script uses the same signature. In case 1 the model answers "awful", which is not one of the three labels, with fallback turned off. In case 2 the model ignores the markers and answers in JSON, with fallback left on.

# parse_error.py
from typing import Literal

import dspy
from dspy.utils import DummyLM
from dspy.utils.exceptions import AdapterParseError


class ReviewLabel(dspy.Signature):
    """Label the sentiment of a product review."""

    review: str = dspy.InputField()
    sentiment: Literal["positive", "negative", "neutral"] = dspy.OutputField()
    stars: int = dspy.OutputField(desc="1 to 5")


classify = dspy.Predict(ReviewLabel)

print("1. A value outside the Literal, fallback off")
lm = DummyLM([{"sentiment": "awful", "stars": 2}])
with dspy.context(lm=lm, adapter=dspy.ChatAdapter(use_json_adapter_fallback=False)):
    try:
        classify(review="Broke after two days.")
    except AdapterParseError as e:
        lines = str(e).splitlines()
        print(lines[0])
        print(lines[1])
        print([l for l in lines if l.startswith("Adapter ")][0])
print("calls made:", len(lm.history))

print()
print("2. The model answers in JSON, ignoring the markers; fallback on (the default)")
good = {"sentiment": "negative", "stars": 2}
json_lm = DummyLM([good, good], adapter=dspy.JSONAdapter())
with dspy.context(lm=json_lm, adapter=dspy.ChatAdapter()):
    pred = classify(review="Broke after two days.")
print("result:", repr(pred.sentiment), repr(pred.stars))
print("calls made:", len(json_lm.history))
for i, h in enumerate(json_lm.history, start=1):
    print(f"call {i} asked:", h["messages"][-1]["content"].split("\n\n")[-1][:60], "...")

Giving DummyLM a JSONAdapter makes the fake model reply in JSON, which is how case 2 stages a model that ignores the markers. Run with tmp/dspy-venv/Scripts/python.exe parse_error.py. Output:

1. A value outside the Literal, fallback off
Failed to parse field sentiment with value awful from the LM response. Error message: 1 validation error for literal['positive','negative','neutral']
  Input should be 'positive', 'negative' or 'neutral' [type=literal_error, input_value='awful', input_type=str]
Adapter ChatAdapter failed to parse the LM response. 
calls made: 1

2. The model answers in JSON, ignoring the markers; fallback on (the default)
result: 'negative' 2
calls made: 2
call 1 asked: Respond with the corresponding output fields, starting with  ...
call 2 asked: Respond with a JSON object in the following order of fields: ...

The error names the field, the bad value, and the allowed values. The full message also carries the raw reply and the list of fields it expected, which is what you read first when debugging. Case 2 succeeded, but it cost two model calls instead of one.

In an optimization pipeline

Adapters belong to stage 1, defining the program, but they matter at every later stage. An optimizer only changes instructions and demos (worked examples); the adapter decides how those land in the messages. When a candidate instruction confuses the model into a broken format, the rollout fails to parse, and the metric sees a failure. dspy.GEPA turns this into a score of failure_score (0.0 by default), and it has an add_format_failure_as_feedback option that shows such failures to the reflection model. If you fine-tune in stage 4, the adapter also formats the training data, so train and serve with the same adapter. In 3.4.0 only ChatAdapter can do this; JSONAdapter's training-data formatter raises NotImplementedError.

Common mistakes

  • Free-text labels. A str output for a label lets the model answer "Negative." or "mostly negative". Your metric then marks correct answers wrong. Use Literal[...].
  • Ignoring the fallback cost. A small local model that often breaks the marker format makes every such call twice. You see it as a call count higher than your example count, and as slow runs. Check with len(lm.history) or the LM's usage tracking, and consider JSONAdapter from the start.
  • Assuming JSON mode guarantees types. JSON mode only promises valid JSON. With JSONAdapter, the field types are still checked after parsing, so an int field holding "two" is still a parse error.
  • Open-ended dict outputs with JSONAdapter. A dict[str, Any] output cannot become a strict schema, so the source falls back to plain JSON mode for that signature. Prefer a Pydantic model with named fields.
  • Swallowing the error. Catching AdapterParseError and returning an empty prediction hides a prompt problem. Let it surface during development; during evaluation, dspy.Evaluate scores it with failure_score.

Cost

Formatting and parsing cost microseconds; the cost is in tokens and calls. The field list and type notes add a fixed block to every prompt, so for a signature with ff fields the overhead grows roughly in proportion to ff, and it is paid on every call. The fallback is the bigger lever: if a fraction qq of replies fail to parse with ChatAdapter, a run of NN calls makes about N(1+q)N(1 + q) calls, and each extra call costs its input tokens again. With cinc_{\text{in}} dollars per input token, coutc_{\text{out}} per output token, and tint_{\text{in}}, toutt_{\text{out}} tokens per call, the run costs about N(1+q)(tincin+toutcout)N(1 + q)(t_{\text{in}} c_{\text{in}} + t_{\text{out}} c_{\text{out}}) dollars. Builder time: picking types well up front is minutes; debugging an untyped label field after an optimization run is hours.

Going further

  • Composing programs: several predictors in one program, each formatted by the same adapter.
  • Writing metrics: what the metric sees when a reply fails to parse.
  • The dspy.XMLAdapter and dspy.TwoStepAdapter entries in the DSPy adapters API reference.
  • The ChatAdapter and JSONAdapter source in your installed DSPy, about 600 lines between them, which is the final word on formats.

Back to DSPy and GEPA: programming and optimizing language-model systems