technique
Saving and serving programs
Keeping an optimized program: saving state versus the whole program, loading it in production, pinning models and versions, and serving it behind an API.
Before this
This page assumes you are comfortable with:
Why you need this
An optimizer run costs hours and real model calls. What it produces is small: a few instructions and worked examples per predictor. If you do not save that, the run is gone when the Python process ends. This page covers stage 5's first job: writing the result to disk, loading it somewhere else, proving it behaves the same, and putting it behind an API.
The idea
A DSPy program is a module tree written in Python; a predictor is one dspy.Predict inside it. Optimizers do not change your Python code. They change each predictor's state: its instruction text and its list of demos (worked examples). So there are two things you could save:
| State only | Whole program | |
|---|---|---|
| Call | program.save("program.json") |
program.save("program_dir", save_program=True) |
| Path | A file ending .json (or .pkl) |
A directory, no file extension |
| Contains | Instructions, demos, field prefixes, each predictor's model settings if set, version metadata | program.pkl (the whole object, pickled with cloudpickle) and metadata.json |
| Load with | Build the same program in code, then program.load("program.json") |
dspy.load("program_dir", allow_pickle=True), no code needed |
| Readable | Yes, by a person and by a diff tool | No |
| Safe to load from others | Yes | No: unpickling can run arbitrary code |
These names and rules come from the DSPy 3.4 source (dspy/primitives/base_module.py and dspy/utils/saving.py). One difference from the DSPy saving tutorial: in 3.4, dspy.load refuses to run unless you pass allow_pickle=True, which the tutorial's example omits.
Use state-only JSON by default. Your program's structure already lives in your code and your version control; the JSON holds exactly what the optimizer learned, and you can read it.
Worked example
This file compiles a tiny short-answer program with dspy.BootstrapFewShot on DSPy's fake model, DummyLM (it returns scripted answers, so no network call is made), saves the state, loads it into a fresh program, and checks that both behave the same. Verified on DSPy 3.4.0 (ran on Python 3.14; 3.12 and newer behave the same).
import dspy
from dspy.utils import DummyLM
trainset = [
dspy.Example(question="What is the capital of France?", answer="Paris").with_inputs("question"),
dspy.Example(question="How many legs does a spider have?", answer="8").with_inputs("question"),
]
def exact_match(example, prediction, trace=None):
return example.answer == prediction.answer
# Compile: the fake teacher answers both training questions correctly,
# so BootstrapFewShot keeps both as demos.
dspy.configure(lm=DummyLM([{"answer": "Paris"}, {"answer": "8"}]))
optimizer = dspy.BootstrapFewShot(metric=exact_match, max_bootstrapped_demos=2, max_labeled_demos=0)
compiled = optimizer.compile(dspy.Predict("question -> answer"), trainset=trainset)
# Save the state only, as JSON.
compiled.save("program.json")
# Load the state into a fresh program built from the same code.
fresh = dspy.Predict("question -> answer")
fresh.load("program.json")
print("same state:", fresh.dump_state() == compiled.dump_state())
# Run both on one new question with identical fake answers.
lm_a = DummyLM([{"answer": "Mercury"}])
lm_b = DummyLM([{"answer": "Mercury"}])
q = "Which planet is closest to the Sun?"
with dspy.context(lm=lm_a):
out_a = compiled(question=q)
with dspy.context(lm=lm_b):
out_b = fresh(question=q)
print("same answer:", out_a.answer == out_b.answer, out_a.answer)
print("same prompt:", lm_a.history[-1]["messages"] == lm_b.history[-1]["messages"])
print("messages sent:", len(lm_a.history[-1]["messages"]))
Run with python save_load.py. Output, with the progress bar trimmed:
...
Bootstrapped 2 full traces after 1 examples for up to 1 rounds, amounting to 2 attempts.
same state: True
same answer: True Mercury
same prompt: True
messages sent: 6
The answer check alone proves little, because the fake model returns "Mercury" whatever it is asked. The prompt check is the real proof: both programs sent the model the exact same six messages (the system message, two demos as user and assistant pairs, and the new question). Same messages to the same model means the same behavior.
This is the whole program.json:
{
"traces": [],
"train": [],
"demos": [
{
"augmented": true,
"question": "What is the capital of France?",
"answer": "Paris"
},
{
"augmented": true,
"question": "How many legs does a spider have?",
"answer": "8"
}
],
"signature": {
"instructions": "Given the fields `question`, produce the fields `answer`.",
"fields": [
{
"prefix": "Question:",
"description": "${question}"
},
{
"prefix": "Answer:",
"description": "${answer}"
}
]
},
"lm": null,
"metadata": {
"dependency_versions": {
"python": "3.14",
"dspy": "3.4.0",
"cloudpickle": "3.1"
}
}
}
How to read it as a person:
signature.instructionsis the instruction the model sees. BootstrapFewShot does not change instructions, so this is DSPy's default. After GEPA or MIPROv2 it holds the optimized instruction, often several paragraphs: read it, because it tells you what the optimizer decided the task is.demosare the worked examples."augmented": truemarks demos the program produced itself during bootstrapping, as opposed to copied labels.lmisnullbecause the model was set globally withdspy.configure. If a predictor has its own model (set_lm), its settings are saved here: model name, temperature, and so on, but never API keys. In 3.4, anapi_basesaved here is dropped at load time with a warning unless you passallow_unsafe_lm_state=True, so a file from elsewhere cannot silently redirect your calls.metadata.dependency_versionsrecords Python, DSPy, and cloudpickle versions. On load, DSPy logs a warning for any mismatch.
A program with several predictors gets one block like this per predictor, keyed by the predictor's name.
The whole-program form
Verified on DSPy 3.4.0, run from the same folder:
import os
import dspy
from dspy.utils import DummyLM
program = dspy.ChainOfThought("question -> answer")
program.save("whole_program", save_program=True)
print(sorted(os.listdir("whole_program")))
try:
dspy.load("whole_program")
except ValueError as err:
print("refused:", err)
loaded = dspy.load("whole_program", allow_pickle=True)
print(type(loaded).__name__)
with dspy.context(lm=DummyLM([{"reasoning": "It orbits nearest.", "answer": "Mercury"}])):
print(loaded(question="Which planet is closest to the Sun?").answer)
Output of python save_whole.py:
... WARNING dspy.primitives.base_module: Loading untrusted .pkl files can run arbitrary code, which may be dangerous. To avoid this, prefer saving using json format using module.save("module.json").
['metadata.json', 'program.pkl']
refused: Loading with pickle is not allowed. Please set `allow_pickle=True` if you are sure you trust the source of the model.
ChainOfThought
Mercury
If your program uses classes from your own modules, pass them as modules_to_serialize=[my_module] so they are pickled by value and the loading machine does not need your source.
Pinning versions
A saved program is only reproducible if everything that shapes the prompt and the answer is pinned:
| Pin | Why |
|---|---|
DSPy version (dspy==3.4.0 in your requirements) |
The adapter turns state into messages. A different DSPy can format the same JSON into different messages. The DSPy docs promise compatibility only within a major release from 3.0 on. |
| Python version | Recorded in the metadata; mainly matters for the pickle form. |
| Task model tag | The local task model is served by Ollama, connected with dspy.LM("ollama_chat/llama3.2", api_base="http://localhost:11434", api_key=""). If the tag is re-pulled after its publisher updates it, your optimized instruction now meets a different model. Record the exact tag and, where your server shows one, the weights' digest. |
| Reflection or teacher model | Claude Opus 5.5 as anthropic/claude-opus-5-5, with its key from the ANTHROPIC_API_KEY environment variable. It matters only when you re-optimize, but record it with the run. |
| Sampling settings | Temperature and maximum tokens, set on the dspy.LM. |
Serving
To serve a program, load it once at startup and call it per request. DSPy 3.4 provides dspy.asyncify(program), which runs a program in a worker thread pool so an async web server is not blocked; dspy.configure(async_max_workers=N) sets the pool size (the DSPy deployment tutorial gives the default as 8). Verified on DSPy 3.4.0, after save_load.py has written program.json:
import asyncio
import dspy
from dspy.utils import DummyLM
dspy.configure(lm=DummyLM([{"answer": "Paris"}, {"answer": "8"}]), async_max_workers=4)
program = dspy.Predict("question -> answer")
program.load("program.json")
async_program = dspy.asyncify(program)
async def main():
results = await asyncio.gather(
async_program(question="What is the capital of France?"),
async_program(question="How many legs does a spider have?"),
)
print(sorted(r.answer for r in results))
asyncio.run(main())
Output of python serve_async.py:
['8', 'Paris']
The deployment tutorial wraps exactly this in a FastAPI endpoint, and shows dspy.streamify(program) for sending output to the client as it is generated. It also shows logging the program to MLflow and serving it with MLflow's model server. In production, remember the DSPy cache: identical requests return the cached answer without a model call, which saves money but means a changed model behind the same tag may not show up until the cache entry is gone. dspy.configure_cache controls it.
In an optimization pipeline
Saving closes stage 4 and opens stage 5. The saved JSON is the artifact you ship, the thing you diff between optimization runs, and the baseline every later run is compared against. Overfitting and re-optimizing depends on it: you cannot tell whether a new run is better without loading the old one and evaluating both on the same testset. A program tuned by BootstrapFinetune also points at a trained model, which must be stored and served alongside the JSON.
Common mistakes
- Loading into different code.
loadmatches state to predictors by name. Renameself.steptoself.renamedin__init__and loading the old file fails withKeyError: 'renamed'. Keep predictor names stable once a program has been optimized. - A file extension on a whole-program path.
save("prog.json", save_program=True)raises "pathmust point to a directory without a suffix". - Loading a pickle you did not make. Unpickling runs code. Use JSON for anything that crosses a trust boundary.
- Version drift between machines. The load warning "There is a mismatch of dspy version" means the prompt may now be formatted differently. Pin and rebuild.
- An unpinned model tag. Scores fall after a routine model update, with no change in your code or JSON.
- Comparing only answers. With a fake or cached model, outputs match even when prompts differ. Compare the messages, as the worked example does.
Cost
Saving and loading cost nothing measurable: the JSON above is 712 bytes, and an optimized instruction of several paragraphs adds only a few kilobytes per predictor. Loading reads one file. Serving costs one model call per predictor per request; for a program with predictors handling requests per second, the model sees calls per second, and with async workers and an average call time of seconds the pool sustains about requests per second before requests queue. Cached repeats cost no model call. The builder's time goes to keeping the JSON, the DSPy version, and the model tag together, one record per release.
Going further
- The DSPy saving and loading tutorial.
- The DSPy deployment tutorial, for FastAPI, streaming, and MLflow serving.
- The
saveandloadmethods in DSPy'sbase_module.py, which are short and worth reading. - Overfitting and re-optimizing, for deciding when the saved program needs replacing.
Leads to
Back to DSPy and GEPA: programming and optimizing language-model systems