technique
Configuring language models
Connecting DSPy to a local task model served by Ollama and to Claude Opus 5.5 for reflection, with settings, contexts, caching, and usage tracking.
Before this
This page assumes you are comfortable with:
- prerequisiteHow language models generate textTokens, the context window, next-token probabilities, and how temperature and sampling turn one prompt into different outputs.
- prerequisitePython essentials for DSPyJust enough Python to read and run every sample in this cluster: classes and inheritance, type hints, keyword arguments, callables, and virtual environments.
Why you need this
A DSPy program says what the task is; something still has to say which language model runs it, with what settings, and how much it is allowed to cost. Optimization makes this matter more than usual: a single GEPA or bootstrapping run calls the model hundreds or thousands of times, so the choice of model, the cache, and the usage numbers decide whether a run is cheap and repeatable or slow and expensive. This page sets up the two-model arrangement used across the cluster. Every sample ran on DSPy 3.4 (3.4.0); the ones that would call a real model were run with DummyLM, DSPy's fake language model, substituted, so no network call was made.
The idea
dspy.LM and the model string
dspy.LM is the object that represents one model with its settings. Its first argument is a model string of the form "provider/model-name". The part before the slash picks the provider, the service or program that hosts the model, and the rest is that provider's name for the model.
Under the hood, DSPy needs code that speaks each provider's API. In DSPy 3.4 there are two such engines: a native one bundled inside DSPy (called lm15 in the source) and LiteLLM, a separate library that knows many providers and is still installed with DSPy. The default, engine="auto", uses the native engine when it supports the provider and the settings, and falls back to LiteLLM otherwise; engine="litellm" forces LiteLLM. Older DSPy documentation says every call goes through LiteLLM; in 3.4 that is only the fallback. Checked offline in the installed source, both model strings below resolve to the native engine.
The local task model
The task model is the one your program calls on every input, many times. In this cluster it is a small open model served by Ollama, a program that downloads open models and serves them on your own machine behind a local web address. DSPy's language-model documentation connects to it like this:
dspy.LM("ollama_chat/llama3.2", api_base="http://localhost:11434", api_key="")
ollama_chat selects Ollama's chat interface, llama3.2 is the model tag used in DSPy's docs (use whatever tag you pulled into Ollama), and api_base is the address where Ollama listens by default. A local server needs no key, so api_key is empty. How fast it runs depends on your hardware and the model's size; the model's page in Ollama's library lists its download size.
The reflection model
The reflection model (called the teacher when it generates examples or training data) is Claude Opus 5.5, a strong hosted model used sparingly:
dspy.LM("anthropic/claude-opus-5-5")
There is no key in the code. DSPy reads it from the ANTHROPIC_API_KEY environment variable, which you set in your shell or your system's secret store. A key written into a script ends up in version control and in every copy of the file.
Why two models
| Task model | Reflection model or teacher | |
|---|---|---|
| Called | On every input, in every evaluation and every optimizer trial | A few times per optimizer step |
| Wants | Cheap per call, fast, private, under your control | The best judgment you can buy |
| Here | A local model through Ollama | Claude Opus 5.5 |
An optimizer such as GEPA runs the cheap model many times to find out what goes wrong, then asks the strong model a few times how to fix the instruction. You pay strong-model prices only for the thinking, and the program you ship runs on the cheap model.
Settings
Settings passed to dspy.LM become defaults for every call it makes:
| Argument | Default in 3.4 | Meaning |
|---|---|---|
temperature |
not set (provider default) | How random sampling is (How language models generate text) |
max_tokens |
not set | Most tokens the model may generate per reply |
cache |
True |
Reuse stored replies for identical requests |
num_retries |
3 |
Retries on temporary failures such as rate limits, with growing waits |
api_base, api_key |
not set | Where to connect and with which key |
dspy.configure versus dspy.context
dspy.configure(lm=...) sets the default model for the whole process, and it can be called again only from the thread that called it first. with dspy.context(lm=...): overrides the setting inside one with block and restores it on exit; it works from any thread. A third option sets the model on one module: module.set_lm(lm) gives every predictor inside it that model. A predictor's own model wins over both configure and context.
The cache
DSPy caches replies by default, in memory and on disk (in .dspy_cache under your home folder unless the DSPY_CACHEDIR environment variable says otherwise, with a 30 GB default limit). The cache key is the whole request (model, messages, and settings) minus the key and the address, so an identical request returns the stored reply with no call and no charge. That makes reruns of an evaluation free and repeatable, and it also means rerunning an unchanged program never shows you sampling noise. To get fresh replies:
dspy.LM(..., cache=False)turns the cache off for that model; a single call can also passcache=False.- A
rollout_idsetting (any integer) becomes part of the cache key, sorollout_id=1androllout_id=2are stored separately. It only produces different replies whentemperatureis above 0, and DSPy warns if you combine it with temperature 0. dspy.configure_cache(enable_disk_cache=False, enable_memory_cache=False)turns caching off for the process.
Usage tracking
dspy.configure(track_usage=True) makes each module call add up the tokens it used, per model. prediction.get_lm_usage() returns them as a dictionary keyed by model name. Replies served from the cache are not counted, because they cost nothing. Each model also keeps lm.history, a list with one entry per call holding the messages, the outputs, the token usage, and a cost when the engine reports one.
Worked example
The configuration lives in its own file, so the program and the experiments import the same models:
# models.py
import dspy
def make_task_lm():
# The local task model, served by Ollama on this machine.
return dspy.LM(
"ollama_chat/llama3.2",
api_base="http://localhost:11434",
api_key="",
temperature=0.0,
max_tokens=1000,
)
def make_reflection_lm():
# The reflection model. The key comes from ANTHROPIC_API_KEY in the environment.
return dspy.LM(
"anthropic/claude-opus-5-5",
temperature=1.0,
max_tokens=32000,
)
Temperature 0 keeps the task model's answers steady for scoring. The reflection settings are the ones the dspy.GEPA source recommends for a reflection model; check the model's documentation for its own output limit.
This script builds the real objects (building one makes no call), then swaps in two fake models with scripted answers and runs a support-ticket triage predictor three times: twice on the default model and once inside a dspy.context block.
# run_fake.py
import dspy
from dspy.utils import DummyLM
from models import make_reflection_lm, make_task_lm
# Build the real configuration objects. Nothing is sent: building an LM makes no call.
real_task, real_reflection = make_task_lm(), make_reflection_lm()
print("task model: ", real_task.model, real_task.kwargs["temperature"])
print("reflection model:", real_reflection.model, real_reflection.kwargs["temperature"])
# Substitute fake LMs with scripted answers so the run makes no network call.
task_lm = DummyLM([{"label": "billing"}, {"label": "shipping"}])
task_lm.model = "fake-task"
reflection_lm = DummyLM([{"label": "billing"}])
reflection_lm.model = "fake-reflection"
dspy.configure(lm=task_lm, track_usage=True)
triage = dspy.Predict("ticket -> label")
a = triage(ticket="I was charged twice this month.")
b = triage(ticket="My parcel has not arrived.")
with dspy.context(lm=reflection_lm):
c = triage(ticket="I was charged twice this month.")
print(a.label, b.label, c.label)
print("calls to task model: ", len(task_lm.history))
print("calls to reflection model:", len(reflection_lm.history))
print("usage of last call:", list(c.get_lm_usage()))
python run_fake.py prints:
task model: ollama_chat/llama3.2 0.0
reflection model: anthropic/claude-opus-5-5 1.0
billing shipping billing
calls to task model: 2
calls to reflection model: 1
usage of last call: ['fake-reflection']
The first two calls went to the default model; the third, inside the with block, went to the other one, and its usage is recorded under that model's name. To run against the real models, delete the four fake-model lines and pass make_task_lm() to dspy.configure; that needs Ollama running and the key set.
Which component uses which model in a typical optimization run:
| Component | Model | Set with |
|---|---|---|
| The program's predictors | Local task model | dspy.configure(lm=make_task_lm()) |
| Evaluation runs | Local task model | Same default |
| GEPA's reflection step | Claude Opus 5.5 | The optimizer's reflection_lm= argument |
| Teacher in bootstrapping or fine-tuning | Claude Opus 5.5 | The optimizer's teacher settings (for dspy.BootstrapFewShot, teacher_settings=dict(lm=...)) |
| An LLM-as-judge metric | Claude Opus 5.5 | dspy.context(lm=...) or set_lm on the judge module |
In an optimization pipeline
This is stage 1, Define the program: the models are part of the program's definition, alongside Signatures and modules. Stage 2 depends on it directly, because cache settings decide whether a rerun measures anything new and usage tracking tells you what an evaluation costs. In stage 3, GEPA in practice passes the reflection model to the optimizer. In stage 5, a pinned model tag is part of what you save, because a different model can need a different optimized prompt.
Common mistakes
- Writing
dspy.configure(lm="ollama_chat/llama3.2")with a bare string. The configure call succeeds, then the first module call raises aValueErrorasking fordspy.LM(...)instead. - Putting the API key in the code. The script works, then the key is in the repository history and has to be revoked.
- Expecting new answers on a rerun. The cache returns identical replies, so a "noise check" shows zero noise. Use
rollout_idwith a temperature above 0, or turn the cache off. - Using
rollout_idat temperature 0. You get a separate cache entry for the same answer, and a warning. - Wrong model tag or Ollama not running. Building the
dspy.LMsucceeds, because it makes no call; the first real call fails with a connection or unknown-model error. Make one test call before starting a long run. - Calling
dspy.configurefrom worker threads. It raises an error saying only the thread that first configured settings may change them; usedspy.contextthere.
Cost
Let and be the number of calls to each model in a run, and the average input and output tokens per call (measured with usage tracking), and and each model's price per token. The hosted model costs dollars. The local model has no per-token price; its cost is wall-clock time, roughly times the average seconds per call divided by the number of calls running at once, plus the machine it runs on. Cache hits cost neither. The disk cache costs disk space, up to its 30 GB default limit.
Going further
- The DSPy Language Models documentation, for other providers and the
responsesmodel type. - Ollama's model library, to choose a task model that fits your machine.
- The DSPy 3.4 migration notes on the normalized LM interface, for the engine choice.
- GEPA in practice, where the reflection model is put to work.
Leads to
Back to DSPy and GEPA: programming and optimizing language-model systems