Topic hub

DSPy and GEPA: programming and optimizing language-model systems

How to write a language-model task as a program instead of a prompt string, measure it, and let an optimizer improve its prompts or its weights, with GEPA's reflective prompt evolution at the center.

The problem

You have a task a language model can almost do: sort support tickets, grade a summary, answer questions from a document. You write a prompt, it works on the five examples you tried, and then it fails on the sixth. You add a sentence to fix that case, and an old case breaks. You switch to a cheaper model, and the whole prompt has to be rewritten. The prompt has become a fragile string that nobody can measure or maintain.

DSPy is a Python framework that replaces that string with a program. You declare what goes in and what comes out (a signature), wire steps together as ordinary Python (modules), and write down what a good answer means (a metric). Then an optimizer does the tuning you were doing by hand: it picks worked examples, rewrites instructions, or even fine-tunes the model's weights, and it checks every change against your metric. GEPA is the optimizer this cluster spends the most time on. It runs the program, reads in plain language why each output was wrong, and asks a strong model to write a better instruction, keeping a varied set of good candidates as it goes.

Underneath are first-year ideas about probability, search, and how models learn, and each of those has its own page. Follow the "Before this" links downward until you land on something you already know, then read back up.

The pipeline

Every page points back to this picture of how a language-model program gets built and improved.

Stage Question it answers Techniques that live here
1. Define the program What is the task, written as signatures and modules instead of prompt strings? Signatures and modules, Composing programs, Adapters and structured output, Configuring language models
2. Measure it What counts as a good output, and how sure is the score? Building a dataset, Writing metrics, LLM-as-judge metrics, Evaluating programs
3. Optimize the prompts How do instructions and examples get better without hand-editing? Few-shot bootstrapping, MIPROv2, GEPA: reflective prompt evolution, GEPA in practice
4. Optimize the weights When is it worth changing the model itself, and how? BootstrapFinetune, Reinforcement learning with GRPO, Prompts and weights together
5. Ship and maintain How does an optimized program get saved, served, and re-optimized when the model or data changes? Saving and serving programs, Overfitting and re-optimizing

Stage 2 is where most projects go wrong. An optimizer climbs whatever metric you give it, so a careless metric produces a program that scores well and behaves badly, and a dev set of twenty examples produces scores that move by ten points from noise alone. Every optimizer page assumes stage 2 was done honestly.

Which page for which job

You have never used DSPy. Read Signatures and modules, then Configuring language models, then Building a dataset and Writing metrics. With those four you can run a program and score it, which is the starting point for everything else.

You have a working program and want it better. Start with Evaluating programs to get a baseline you trust, then Few-shot bootstrapping (cheap, often enough), then GEPA. Read the algorithm page before GEPA in practice; the settings make sense only once you know what each one controls.

You want to understand GEPA itself. The algorithm page leans on three prerequisites: Evolutionary algorithms for the population and mutation idea, Pareto fronts for how candidates are kept, and Writing metrics for the feedback text that makes reflection work.

You run the program thousands of times a day on a small local model. That is the case for stage 4. Read BootstrapFinetune first, then Prompts and weights together, which helps you decide which lever to pull first.

Your scores changed and you do not know why. Overfitting and re-optimizing, then Noisy scores and sample size. Often the program did not get worse; the measurement was never that precise.

The basics underneath

None of these needs anything beyond first-year college material, and most need less.

You need For
How language models generate text Tokens, the context window, and why the same prompt gives different answers.
Prompts and few-shot learning Instructions and worked examples: the two things prompt optimizers change.
Python essentials for DSPy Classes, type hints, and passing functions around, enough to read every sample.
Structured output and JSON Why programs need model output in typed fields, and what a parse failure looks like.
Training, validation, and test splits Which examples an optimizer may learn from, choose with, and report on.
Noisy scores and sample size How much a score wobbles, and how many examples make a difference real.
Search and optimization basics Objectives, budgets, hill climbing, and exploring versus exploiting.
Evolutionary algorithms Populations, selection, mutation, and merge: the family GEPA belongs to.
Pareto fronts Keeping candidates that are each best at something different.
Gradient descent and fine-tuning How weights change during training, and what fine-tuning changes.
LoRA and parameter-efficient tuning Training a small add-on instead of the whole model.
Reinforcement learning basics Learning from a reward instead of a correct answer.

Conventions used across this cluster

So the pages agree with each other:

  • Code is Python with DSPy 3.4. Samples that can run were run against DSPy's fake language model, which returns scripted answers, so their output is real but the answers are staged. Samples that need a real model or training hardware say so in the sentence before them, and any result they show is an illustration.
  • Two models. The task model is a small local model served by Ollama, run many times. The reflection model (for GEPA) and the teacher (for bootstrapping and fine-tuning) is Claude Opus 5.5, used sparingly. The pages explain why this split is the usual one. One exception: DSPy cannot fine-tune a model served by Ollama, so the weight pages in stage 4 train a Hugging Face copy of a small model instead and say so.
  • Scores always carry their split and size. "62% on a 50-example valset", never a bare "62%".
  • Words with fixed meanings: a program is a DSPy module tree, a predictor is one model call inside it, a candidate is one version of the program during optimization, a rollout is one run on one example, and a trace is the record of every predictor's inputs and outputs during a rollout.
  • Hardware is described generically, as "a GPU with N GB of memory", with numbers from the published documentation.
  • There is no shared running example yet; each page uses its own small task.

Names only, no links: names stay stable, download pages do not.

Job Pick Why
The framework DSPy Signatures, modules, metrics, evaluation, and every optimizer in this cluster.
GEPA outside DSPy gepa (the standalone package) The same optimizer for any text you can score: a system prompt, a config, a piece of code. DSPy installs it for you.
Running the local task model Ollama Serves open models on your own machine behind a simple API that DSPy can call.
Seeing what the program did MLflow tracing, or DSPy's own inspect_history Every prompt sent and every answer received, which is how you debug a metric or a module.
Writing and running experiments Any Python editor with a terminal, or a notebook Optimization runs are long; a notebook keeps results around between steps.

Prices and versions are deliberately absent, apart from the DSPy version the samples were checked against.

How to read this cluster

The learning path below is sorted so that each row depends only on the rows above it. If you already know how language models work and have written Python, start at Signatures and modules. If you came for GEPA, start at its algorithm page and follow its links back only as far as you need.

Learning path

Each row depends only on rows above it. Read top to bottom, or jump to a technique and follow its "Before this" links downward.