technique

Building a dataset

Turning examples into dspy.Example objects, marking which fields are inputs, deciding how many you need, and splitting them for optimization.

Before this

This page assumes you are comfortable with:

Why you need this

No optimizer can improve a program it cannot measure, and measuring needs examples. In DSPy the examples are a plain Python list of dspy.Example objects. Building that list well, with the right fields marked as inputs, the hard cases included, and a fixed split, is most of stage 2. It is also the part nobody can do for you, because only you know what the task's real inputs look like.

The idea

A dspy.Example is one row of data with named fields, like a dictionary you can also read with dots: ex.question, ex.answer. On its own it does not know which fields go into the program. You say so with with_inputs:

ex = dspy.Example(question="What is the capital of Canada?", answer="Ottawa").with_inputs("question")

After that, two methods split the row:

Method Returns Used by
ex.inputs() The fields named in with_inputs, here question The program: program(**ex.inputs())
ex.labels() Every other field, here answer The metric, which compares them with the program's output

with_inputs returns a new example and leaves the original unchanged, so write it at the end of the line that builds the example.

Labels versus inputs. An input is something the program will have in real use. A label is something only you know: the correct answer, a category, a note for the metric. Extra fields are fine. A kind field saying which topic a question is about never reaches the program, and it lets you stratify the split and group failures later.

Label-free examples. Some metrics need no gold answer: "is the summary under 50 words", "does the reply avoid promising a refund". Then an example can be inputs only, dspy.Example(text=...).with_inputs("text"), and labels() is empty. The DSPy metrics guide calls these label-free examples.

How many

DSPy works with small datasets. Exact numbers per optimizer are not given in the DSPy optimizer documentation; these are the facts from the DSPy 3.4.0 source:

Optimizer What it does with your sets
dspy.LabeledFewShot Copies up to k examples (16 by default) from the trainset into the prompt as demos
dspy.BootstrapFewShot Runs the trainset, keeps up to 4 passing traces (max_bootstrapped_demos=4) plus up to 16 labeled demos (max_labeled_demos=16) per predictor
dspy.MIPROv2 Needs a valset of at least 1 example; with none, it moves 80% of your trainset into a valset
dspy.GEPA Trainset for reflection, valset for its per-example Pareto scores; with no valset it reuses the trainset, and it suggests a smaller valset once yours has more than 35 examples

A practical reading: 30 examples is enough to start; a few hundred makes scores precise enough to compare programs (see Noisy scores and sample size).

Why GEPA wants its own valset. GEPA keeps every candidate that is best on at least one validation example. Those per-example scores are how it decides which candidates stay. If the valset is the trainset it learns from, it keeps candidates that fit those exact examples, which the GEPA log message itself warns makes it overfit.

Hard examples on purpose

Random real inputs are mostly easy. An optimizer learns most from failures, so add the cases you know are hard: ambiguous wording, answers that need a unit, inputs in the wrong language, the edge case a user complained about. Mark them with a field (kind="hard") so you can see whether they improved.

Worked example

Thirty short-answer questions in three topics, ten each, read from CSV text, turned into examples, and split 12 / 9 / 9 with every topic equally represented in every set. In a real project the CSV is a file; here it is inline so the script runs on its own. Run on DSPy 3.4.0 (on Python 3.14; 3.12 and newer behave the same).

# build_dataset.py
import csv
import io
import random

import dspy

# 30 short-answer questions. In a real project this text lives in qa.csv.
CSV_TEXT = """question,answer,kind
What is the capital of France?,Paris,geography
What is the capital of Japan?,Tokyo,geography
What is the capital of Canada?,Ottawa,geography
What is the capital of Australia?,Canberra,geography
Which river flows through Cairo?,Nile,geography
What is the largest ocean on Earth?,Pacific,geography
Which continent is Kenya in?,Africa,geography
What is the tallest mountain on Earth?,Everest,geography
What is the longest river in South America?,Amazon,geography
Which country is the city of Lagos in?,Nigeria,geography
How many sides does a hexagon have?,6,math
What is 12 times 12?,144,math
What is the square root of 81?,9,math
How many degrees are in a right angle?,90,math
What is 15 percent of 200?,30,math
How many minutes are in 3 hours?,180,math
What is the next prime after 7?,11,math
How many edges does a cube have?,12,math
How many centimeters are in a meter?,100,math
What is 7 plus 8?,15,math
What gas do plants take in for photosynthesis?,carbon dioxide,science
What is the chemical symbol for gold?,Au,science
How many planets orbit the Sun?,8,science
What is H2O commonly called?,water,science
What part of the cell holds the DNA?,nucleus,science
At what Celsius temperature does water boil at sea level?,100,science
What force keeps the Moon in orbit?,gravity,science
Which planet is known as the Red Planet?,Mars,science
What is the hardest natural mineral?,diamond,science
How many bones are in the adult human body?,206,science
"""

rows = list(csv.DictReader(io.StringIO(CSV_TEXT)))
examples = [
    dspy.Example(question=r["question"], answer=r["answer"], kind=r["kind"]).with_inputs("question")
    for r in rows
]
print("examples:", len(examples))

first = examples[0]
print("inputs:", first.inputs())
print("labels:", first.labels())

# Stratified split: shuffle each kind with a fixed seed, then 4 / 3 / 3 per kind.
rng = random.Random(0)
trainset, valset, testset = [], [], []
for kind in ["geography", "math", "science"]:
    group = [e for e in examples if e.kind == kind]
    rng.shuffle(group)
    trainset += group[:4]
    valset += group[4:7]
    testset += group[7:]

for name, split in [("trainset", trainset), ("valset", valset), ("testset", testset)]:
    counts = {k: sum(e.kind == k for e in split) for k in ["geography", "math", "science"]}
    print(f"{name}: {len(split)} {counts}")

print("valset questions:")
for e in valset:
    print(" ", e.question, "->", e.answer)

Run with tmp/dspy-venv/Scripts/python.exe build_dataset.py:

examples: 30
inputs: Example({'question': 'What is the capital of France?'}) (input_keys={'question'})
labels: Example({'answer': 'Paris', 'kind': 'geography'}) (input_keys=None)
trainset: 12 {'geography': 4, 'math': 4, 'science': 4}
valset: 9 {'geography': 3, 'math': 3, 'science': 3}
testset: 9 {'geography': 3, 'math': 3, 'science': 3}
valset questions:
  What is the capital of Australia? -> Canberra
  Which river flows through Cairo? -> Nile
  What is the capital of Canada? -> Ottawa
  How many sides does a hexagon have? -> 6
  What is 12 times 12? -> 144
  How many edges does a cube have? -> 12
  How many planets orbit the Sun? -> 8
  At what Celsius temperature does water boil at sea level? -> 100
  What force keeps the Moon in orbit? -> gravity

Step by step:

  1. csv.DictReader turns each line into a dictionary keyed by the header row.
  2. Each dictionary becomes a dspy.Example, with only question marked as input. answer and kind are labels.
  3. The split shuffles each topic separately with random.Random(0), a seeded generator, so a rerun gives exactly the same sets. Each topic's ten examples go 4 / 3 / 3.
  4. The counts confirm the stratification: every set has the same topic mix.

DataLoader (imported with from dspy.datasets import DataLoader) also has from_csv, but in 3.4.0 it relies on the separate datasets package, which this venv does not have; the standard csv module needs nothing extra.

A second script shows the two ways examples meet a program and a metric, with DummyLM, DSPy's fake LM that returns scripted answers. In real use the local task model, a model served by Ollama, answers instead.

# use_examples.py
import dspy
from dspy.utils import DummyLM

gold = dspy.Example(question="What is the capital of Canada?", answer="Ottawa")

try:
    gold.inputs()
except ValueError as e:
    print("ValueError:", e)

gold = gold.with_inputs("question")
dspy.configure(lm=DummyLM([{"answer": "Ottawa"}]))
qa = dspy.Predict("question -> answer")
pred = qa(**gold.inputs())
print("pred.answer:", pred.answer, "| gold.answer:", gold.answer)

unlabeled = dspy.Example(text="The checkout page froze twice.").with_inputs("text")
print("label-free labels():", unlabeled.labels())

Output:

ValueError: Inputs have not been set for this example. Use `example.with_inputs()` to set them.
pred.answer: Ottawa | gold.answer: Ottawa
label-free labels(): Example({}) (input_keys=None)

In an optimization pipeline

This is stage 2, measuring the program. The three lists go straight into stage 3: optimizer.compile(program, trainset=trainset, valset=valset) for the optimizers that take a valset, and dspy.Evaluate(devset=valset, metric=...) for a baseline. The testset waits until the end. Save the three lists (as JSON or CSV) next to the optimization run so the run can be repeated and compared later.

Common mistakes

  • Forgetting with_inputs. ex.inputs() raises the ValueError shown above, which usually surfaces deep inside an optimizer's first call.
  • Marking a label as an input. with_inputs("question", "answer") hands the answer to the program. Symptom: a near-perfect score from the first run.
  • Field names that do not match the signature. An example with query for a signature that expects question does not raise an error in 3.4.0. dspy.Predict logs two warnings ("Input contains fields not in signature" and "Not all input fields were provided") and calls the model with the question missing. Symptom: confident nonsense and a low score. Rename the fields once, when you build the examples.
  • Only easy examples. The baseline scores 95% and the optimizer has nothing to learn from. Add hard cases until the baseline fails on some.
  • An unseeded shuffle. Each run gets different sets, so scores from different days cannot be compared.

Cost

Building the list is linear in the number of examples nn and takes milliseconds. The cost is the builder's time: writing and checking a gold answer takes minutes per example, so 30 examples is an afternoon and 300 is a week of part-time work. Model costs follow the sets: every optimizer evaluation runs the program on some of them, so a set twice as large roughly doubles calls, tokens, and dollars for each evaluation (see Writing metrics and the optimizer pages for per-optimizer counts). File size is negligible: the 30 questions above are 1,468 bytes of CSV.

Going further

Leads to

Back to DSPy and GEPA: programming and optimizing language-model systems