prerequisite

Structured output and JSON

Why programs need model output in fields instead of free text, what JSON looks like, and what happens when a model's output does not parse.

Before this

This page assumes you are comfortable with:

Why you need this

A language model returns text, but the code around it needs values: a label to compare, a number to sort by, a true or false to branch on. A DSPy program passes one step's output fields to the next step and to a metric, so every reply has to be turned into typed fields, and some replies will not cooperate. This page covers the format most systems use for that, JSON, and what failure looks like, without any DSPy code.

The idea

Structured output means a model reply laid out as named fields with known types, instead of free prose. The code that reads it is a parser: it turns text into values, or reports that it cannot.

JSON

JSON (JavaScript Object Notation) is a plain-text format for data that nearly every language can read. It has six kinds of value:

Kind Example Python value after parsing
String "billing" str
Number 2, 0.75 int or float
Boolean true, false bool (True, False)
Null null None
Array ["late", "damaged"] list
Object {"urgent": true} dict

An object is a set of key and value pairs in braces; keys are always strings in double quotes. An array is an ordered list in square brackets. Values nest, so an object can hold arrays of objects. The rules are strict: double quotes only, lowercase true, no trailing comma, no comments, and nothing before or after the top-level value.

A schema is a promise about shape

A schema is a description of which fields a piece of data must have and what type each one is. JSON Schema is a standard way to write a schema as JSON itself. For a support-ticket triage reply with three fields:

{
  "properties": {
    "category": {"enum": ["billing", "shipping", "account"], "type": "string"},
    "urgent": {"type": "boolean"},
    "priority": {"type": "integer"}
  },
  "required": ["category", "urgent", "priority"],
  "type": "object"
}

(This is the schema the Python library Pydantic generates for the class in the worked example below, with its title entries removed.) An enum is a fixed list of allowed values.

Parsing has two steps

  1. Syntax. Is the text valid JSON at all? Python's json.loads either returns a value or raises json.JSONDecodeError, naming the line and column where it gave up.
  2. Shape. Does the parsed value match the schema: every required field present, every field the right type, every enum value allowed? This is called validation. A validation library reports which field failed and why.

A reply can pass step 1 and fail step 2, as output C below does.

Typed fields

The types you will see in DSPy signatures map onto JSON like this:

Python type hint JSON Valid Invalid
int integer 2 2.5, "two"
bool boolean true "very"
list[str] array of strings ["late"] "late"
Literal["billing", "shipping"] string with enum "billing" "Billing"

Two ways to get fields from a model

Ask for them in the prompt. Describe the fields and the format in words, then parse whatever comes back. This works with any model, but nothing stops the model from adding a friendly sentence, using a capital letter, or inventing a label.

Use the provider's structured-output mode. Some model providers accept a JSON Schema alongside the prompt and promise a reply that matches it. Parse failures become much rarer, but the feature exists only on some providers and models, each supports a different subset of JSON Schema, and a reply can still match the schema while being wrong (a valid "shipping" for a billing ticket). Check the provider's documentation for what it guarantees.

Worked example

Same request, three model outputs, one parser. The parser is plain Python plus Pydantic, a validation library that DSPy itself depends on (so it is already installed wherever DSPy is). A Pydantic model class declares fields with type hints, and model_validate checks a parsed value against them.

# parse_three.py
import json
from typing import Literal

from pydantic import BaseModel, ValidationError


class Triage(BaseModel):
    category: Literal["billing", "shipping", "account"]
    urgent: bool
    priority: int


outputs = {
    "A": '{"category": "billing", "urgent": true, "priority": 2}',
    "B": 'Sure! Here is the JSON:\n{"category": "billing", "urgent": true, "priority": 2}',
    "C": '{"category": "billing", "urgent": "very", "priority": 2}',
}

for name, text in outputs.items():
    try:
        data = json.loads(text)
    except json.JSONDecodeError as e:
        print(name, "parse error:", e)
        continue
    try:
        ticket = Triage.model_validate(data)
        print(name, "ok:", ticket)
    except ValidationError as e:
        first = e.errors()[0]
        print(name, "schema error:", first["loc"], first["msg"], repr(first["input"]))

python parse_three.py prints:

A ok: category='billing' urgent=True priority=2
B parse error: Expecting value: line 1 column 1 (char 0)
C schema error: ('urgent',) Input should be a valid boolean, unable to interpret input 'very'
Output Step 1: syntax Step 2: shape Why
A passes passes Valid JSON, every field present with the right type
B fails not reached Sure! is not a JSON value, so the parser stops at line 1, column 1, character 0
C passes fails on urgent "very" is a string, and the schema wants a boolean

Output B holds perfectly good data behind a polite sentence. A stricter prompt, a provider's structured-output mode, or a lenient parser that cuts out the text between the first { and the last } would rescue it. Output C cannot be rescued by parsing: the model did not decide whether the ticket is urgent, so the only honest fixes are to ask again or to change the prompt.

Validation libraries also bend a little. In a separate run, Pydantic accepted "urgent": "true" and "priority": "2" (strings) as True and 2, but rejected "Billing" (not in the list), 2.5 for priority ("got a number with a fractional part"), and a reply missing priority altogether ("Field required"). Know which conversions your parser makes before you trust its output.

In an optimization pipeline

In stage 1, Define the program, a signature's output fields and their type hints become the schema, and DSPy's adapters handle the prompt wording and the parsing (Adapters and structured output). In stage 2, Measure it, a reply that fails to parse usually scores zero, so parse failures show up as lost points. In stage 3, Optimize the prompts, an optimizer can improve the instruction until a weaker model stops adding stray prose, which is one of the cheapest wins available.

Common mistakes

  • Parsing with string tricks instead of a parser. Splitting on commas works until a value contains a comma; then fields shift and the wrong value lands in the wrong field.
  • No validation after parsing. The JSON parses, urgent is the string "very", and a later if ticket["urgent"]: is true for every non-empty string.
  • Free-text labels. Without an enum, the model returns Billing, billing issue, and payment, and exact-match scores collapse.
  • Asking for JSON and for reasoning in the same free-text reply. The reasoning sentence comes first and output B's parse error appears on every call.
  • Assuming structured-output mode means correct output. Every reply parses, the score is still low, and nothing in the logs says why.

Cost

Parsing and validation cost microseconds; the cost is in failures. If a fraction ff of replies fail to parse and each failure is retried once, the expected number of model calls per input is 1+f1 + f, and each retry pays the full prompt again. Field names, descriptions, and format rules in the prompt add input tokens to every call. Structured-output modes may restrict which models you can use, which can push you to a more expensive one.

Going further

  • The JSON specification (ECMA-404), which fits on a few pages.
  • The JSON Schema documentation, for enum, required, nested objects, and arrays.
  • Pydantic's documentation on strict and lax mode, which controls the conversions shown above.
  • Adapters and structured output, for how DSPy formats fields and recovers from parse failures.

Leads to

Back to DSPy and GEPA: programming and optimizing language-model systems