technique

Configuring language models

Connecting DSPy to a local task model served by Ollama and to Claude Opus 5.5 for reflection, with settings, contexts, caching, and usage tracking.

Before this

This page assumes you are comfortable with:

Why you need this

A DSPy program says what the task is; something still has to say which language model runs it, with what settings, and how much it is allowed to cost. Optimization makes this matter more than usual: a single GEPA or bootstrapping run calls the model hundreds or thousands of times, so the choice of model, the cache, and the usage numbers decide whether a run is cheap and repeatable or slow and expensive. This page sets up the two-model arrangement used across the cluster. Every sample ran on DSPy 3.4 (3.4.0); the ones that would call a real model were run with DummyLM, DSPy's fake language model, substituted, so no network call was made.

The idea

dspy.LM and the model string

dspy.LM is the object that represents one model with its settings. Its first argument is a model string of the form "provider/model-name". The part before the slash picks the provider, the service or program that hosts the model, and the rest is that provider's name for the model.

Under the hood, DSPy needs code that speaks each provider's API. In DSPy 3.4 there are two such engines: a native one bundled inside DSPy (called lm15 in the source) and LiteLLM, a separate library that knows many providers and is still installed with DSPy. The default, engine="auto", uses the native engine when it supports the provider and the settings, and falls back to LiteLLM otherwise; engine="litellm" forces LiteLLM. Older DSPy documentation says every call goes through LiteLLM; in 3.4 that is only the fallback. Checked offline in the installed source, both model strings below resolve to the native engine.

The local task model

The task model is the one your program calls on every input, many times. In this cluster it is a small open model served by Ollama, a program that downloads open models and serves them on your own machine behind a local web address. DSPy's language-model documentation connects to it like this:

dspy.LM("ollama_chat/llama3.2", api_base="http://localhost:11434", api_key="")

ollama_chat selects Ollama's chat interface, llama3.2 is the model tag used in DSPy's docs (use whatever tag you pulled into Ollama), and api_base is the address where Ollama listens by default. A local server needs no key, so api_key is empty. How fast it runs depends on your hardware and the model's size; the model's page in Ollama's library lists its download size.

The reflection model

The reflection model (called the teacher when it generates examples or training data) is Claude Opus 5.5, a strong hosted model used sparingly:

dspy.LM("anthropic/claude-opus-5-5")

There is no key in the code. DSPy reads it from the ANTHROPIC_API_KEY environment variable, which you set in your shell or your system's secret store. A key written into a script ends up in version control and in every copy of the file.

Why two models

Task model Reflection model or teacher
Called On every input, in every evaluation and every optimizer trial A few times per optimizer step
Wants Cheap per call, fast, private, under your control The best judgment you can buy
Here A local model through Ollama Claude Opus 5.5

An optimizer such as GEPA runs the cheap model many times to find out what goes wrong, then asks the strong model a few times how to fix the instruction. You pay strong-model prices only for the thinking, and the program you ship runs on the cheap model.

Settings

Settings passed to dspy.LM become defaults for every call it makes:

Argument Default in 3.4 Meaning
temperature not set (provider default) How random sampling is (How language models generate text)
max_tokens not set Most tokens the model may generate per reply
cache True Reuse stored replies for identical requests
num_retries 3 Retries on temporary failures such as rate limits, with growing waits
api_base, api_key not set Where to connect and with which key

dspy.configure versus dspy.context

dspy.configure(lm=...) sets the default model for the whole process, and it can be called again only from the thread that called it first. with dspy.context(lm=...): overrides the setting inside one with block and restores it on exit; it works from any thread. A third option sets the model on one module: module.set_lm(lm) gives every predictor inside it that model. A predictor's own model wins over both configure and context.

The cache

DSPy caches replies by default, in memory and on disk (in .dspy_cache under your home folder unless the DSPY_CACHEDIR environment variable says otherwise, with a 30 GB default limit). The cache key is the whole request (model, messages, and settings) minus the key and the address, so an identical request returns the stored reply with no call and no charge. That makes reruns of an evaluation free and repeatable, and it also means rerunning an unchanged program never shows you sampling noise. To get fresh replies:

  • dspy.LM(..., cache=False) turns the cache off for that model; a single call can also pass cache=False.
  • A rollout_id setting (any integer) becomes part of the cache key, so rollout_id=1 and rollout_id=2 are stored separately. It only produces different replies when temperature is above 0, and DSPy warns if you combine it with temperature 0.
  • dspy.configure_cache(enable_disk_cache=False, enable_memory_cache=False) turns caching off for the process.

Usage tracking

dspy.configure(track_usage=True) makes each module call add up the tokens it used, per model. prediction.get_lm_usage() returns them as a dictionary keyed by model name. Replies served from the cache are not counted, because they cost nothing. Each model also keeps lm.history, a list with one entry per call holding the messages, the outputs, the token usage, and a cost when the engine reports one.

Worked example

The configuration lives in its own file, so the program and the experiments import the same models:

# models.py
import dspy


def make_task_lm():
    # The local task model, served by Ollama on this machine.
    return dspy.LM(
        "ollama_chat/llama3.2",
        api_base="http://localhost:11434",
        api_key="",
        temperature=0.0,
        max_tokens=1000,
    )


def make_reflection_lm():
    # The reflection model. The key comes from ANTHROPIC_API_KEY in the environment.
    return dspy.LM(
        "anthropic/claude-opus-5-5",
        temperature=1.0,
        max_tokens=32000,
    )

Temperature 0 keeps the task model's answers steady for scoring. The reflection settings are the ones the dspy.GEPA source recommends for a reflection model; check the model's documentation for its own output limit.

This script builds the real objects (building one makes no call), then swaps in two fake models with scripted answers and runs a support-ticket triage predictor three times: twice on the default model and once inside a dspy.context block.

# run_fake.py
import dspy
from dspy.utils import DummyLM

from models import make_reflection_lm, make_task_lm

# Build the real configuration objects. Nothing is sent: building an LM makes no call.
real_task, real_reflection = make_task_lm(), make_reflection_lm()
print("task model:      ", real_task.model, real_task.kwargs["temperature"])
print("reflection model:", real_reflection.model, real_reflection.kwargs["temperature"])

# Substitute fake LMs with scripted answers so the run makes no network call.
task_lm = DummyLM([{"label": "billing"}, {"label": "shipping"}])
task_lm.model = "fake-task"
reflection_lm = DummyLM([{"label": "billing"}])
reflection_lm.model = "fake-reflection"

dspy.configure(lm=task_lm, track_usage=True)
triage = dspy.Predict("ticket -> label")

a = triage(ticket="I was charged twice this month.")
b = triage(ticket="My parcel has not arrived.")
with dspy.context(lm=reflection_lm):
    c = triage(ticket="I was charged twice this month.")

print(a.label, b.label, c.label)
print("calls to task model:      ", len(task_lm.history))
print("calls to reflection model:", len(reflection_lm.history))
print("usage of last call:", list(c.get_lm_usage()))

python run_fake.py prints:

task model:       ollama_chat/llama3.2 0.0
reflection model: anthropic/claude-opus-5-5 1.0
billing shipping billing
calls to task model:       2
calls to reflection model: 1
usage of last call: ['fake-reflection']

The first two calls went to the default model; the third, inside the with block, went to the other one, and its usage is recorded under that model's name. To run against the real models, delete the four fake-model lines and pass make_task_lm() to dspy.configure; that needs Ollama running and the key set.

Which component uses which model in a typical optimization run:

Component Model Set with
The program's predictors Local task model dspy.configure(lm=make_task_lm())
Evaluation runs Local task model Same default
GEPA's reflection step Claude Opus 5.5 The optimizer's reflection_lm= argument
Teacher in bootstrapping or fine-tuning Claude Opus 5.5 The optimizer's teacher settings (for dspy.BootstrapFewShot, teacher_settings=dict(lm=...))
An LLM-as-judge metric Claude Opus 5.5 dspy.context(lm=...) or set_lm on the judge module

In an optimization pipeline

This is stage 1, Define the program: the models are part of the program's definition, alongside Signatures and modules. Stage 2 depends on it directly, because cache settings decide whether a rerun measures anything new and usage tracking tells you what an evaluation costs. In stage 3, GEPA in practice passes the reflection model to the optimizer. In stage 5, a pinned model tag is part of what you save, because a different model can need a different optimized prompt.

Common mistakes

  • Writing dspy.configure(lm="ollama_chat/llama3.2") with a bare string. The configure call succeeds, then the first module call raises a ValueError asking for dspy.LM(...) instead.
  • Putting the API key in the code. The script works, then the key is in the repository history and has to be revoked.
  • Expecting new answers on a rerun. The cache returns identical replies, so a "noise check" shows zero noise. Use rollout_id with a temperature above 0, or turn the cache off.
  • Using rollout_id at temperature 0. You get a separate cache entry for the same answer, and a warning.
  • Wrong model tag or Ollama not running. Building the dspy.LM succeeds, because it makes no call; the first real call fails with a connection or unknown-model error. Make one test call before starting a long run.
  • Calling dspy.configure from worker threads. It raises an error saying only the thread that first configured settings may change them; use dspy.context there.

Cost

Let NtaskN_{\text{task}} and NreflN_{\text{refl}} be the number of calls to each model in a run, tint_{\text{in}} and toutt_{\text{out}} the average input and output tokens per call (measured with usage tracking), and cinc_{\text{in}} and coutc_{\text{out}} each model's price per token. The hosted model costs Nrefl⋅(tincin+toutcout)N_{\text{refl}} \cdot (t_{\text{in}} c_{\text{in}} + t_{\text{out}} c_{\text{out}}) dollars. The local model has no per-token price; its cost is wall-clock time, roughly NtaskN_{\text{task}} times the average seconds per call divided by the number of calls running at once, plus the machine it runs on. Cache hits cost neither. The disk cache costs disk space, up to its 30 GB default limit.

Going further

  • The DSPy Language Models documentation, for other providers and the responses model type.
  • Ollama's model library, to choose a task model that fits your machine.
  • The DSPy 3.4 migration notes on the normalized LM interface, for the engine choice.
  • GEPA in practice, where the reflection model is put to work.

Leads to

Back to DSPy and GEPA: programming and optimizing language-model systems