Topic hub
DSPy and GEPA: programming and optimizing language-model systems
How to write a language-model task as a program instead of a prompt string, measure it, and let an optimizer improve its prompts or its weights, with GEPA's reflective prompt evolution at the center.
The problem
You have a task a language model can almost do: sort support tickets, grade a summary, answer questions from a document. You write a prompt, it works on the five examples you tried, and then it fails on the sixth. You add a sentence to fix that case, and an old case breaks. You switch to a cheaper model, and the whole prompt has to be rewritten. The prompt has become a fragile string that nobody can measure or maintain.
DSPy is a Python framework that replaces that string with a program. You declare what goes in and what comes out (a signature), wire steps together as ordinary Python (modules), and write down what a good answer means (a metric). Then an optimizer does the tuning you were doing by hand: it picks worked examples, rewrites instructions, or even fine-tunes the model's weights, and it checks every change against your metric. GEPA is the optimizer this cluster spends the most time on. It runs the program, reads in plain language why each output was wrong, and asks a strong model to write a better instruction, keeping a varied set of good candidates as it goes.
Underneath are first-year ideas about probability, search, and how models learn, and each of those has its own page. Follow the "Before this" links downward until you land on something you already know, then read back up.
The pipeline
Every page points back to this picture of how a language-model program gets built and improved.
| Stage | Question it answers | Techniques that live here |
|---|---|---|
| 1. Define the program | What is the task, written as signatures and modules instead of prompt strings? | Signatures and modules, Composing programs, Adapters and structured output, Configuring language models |
| 2. Measure it | What counts as a good output, and how sure is the score? | Building a dataset, Writing metrics, LLM-as-judge metrics, Evaluating programs |
| 3. Optimize the prompts | How do instructions and examples get better without hand-editing? | Few-shot bootstrapping, MIPROv2, GEPA: reflective prompt evolution, GEPA in practice |
| 4. Optimize the weights | When is it worth changing the model itself, and how? | BootstrapFinetune, Reinforcement learning with GRPO, Prompts and weights together |
| 5. Ship and maintain | How does an optimized program get saved, served, and re-optimized when the model or data changes? | Saving and serving programs, Overfitting and re-optimizing |
Stage 2 is where most projects go wrong. An optimizer climbs whatever metric you give it, so a careless metric produces a program that scores well and behaves badly, and a dev set of twenty examples produces scores that move by ten points from noise alone. Every optimizer page assumes stage 2 was done honestly.
Which page for which job
You have never used DSPy. Read Signatures and modules, then Configuring language models, then Building a dataset and Writing metrics. With those four you can run a program and score it, which is the starting point for everything else.
You have a working program and want it better. Start with Evaluating programs to get a baseline you trust, then Few-shot bootstrapping (cheap, often enough), then GEPA. Read the algorithm page before GEPA in practice; the settings make sense only once you know what each one controls.
You want to understand GEPA itself. The algorithm page leans on three prerequisites: Evolutionary algorithms for the population and mutation idea, Pareto fronts for how candidates are kept, and Writing metrics for the feedback text that makes reflection work.
You run the program thousands of times a day on a small local model. That is the case for stage 4. Read BootstrapFinetune first, then Prompts and weights together, which helps you decide which lever to pull first.
Your scores changed and you do not know why. Overfitting and re-optimizing, then Noisy scores and sample size. Often the program did not get worse; the measurement was never that precise.
The basics underneath
None of these needs anything beyond first-year college material, and most need less.
| You need | For |
|---|---|
| How language models generate text | Tokens, the context window, and why the same prompt gives different answers. |
| Prompts and few-shot learning | Instructions and worked examples: the two things prompt optimizers change. |
| Python essentials for DSPy | Classes, type hints, and passing functions around, enough to read every sample. |
| Structured output and JSON | Why programs need model output in typed fields, and what a parse failure looks like. |
| Training, validation, and test splits | Which examples an optimizer may learn from, choose with, and report on. |
| Noisy scores and sample size | How much a score wobbles, and how many examples make a difference real. |
| Search and optimization basics | Objectives, budgets, hill climbing, and exploring versus exploiting. |
| Evolutionary algorithms | Populations, selection, mutation, and merge: the family GEPA belongs to. |
| Pareto fronts | Keeping candidates that are each best at something different. |
| Gradient descent and fine-tuning | How weights change during training, and what fine-tuning changes. |
| LoRA and parameter-efficient tuning | Training a small add-on instead of the whole model. |
| Reinforcement learning basics | Learning from a reward instead of a correct answer. |
Conventions used across this cluster
So the pages agree with each other:
- Code is Python with DSPy 3.4. Samples that can run were run against DSPy's fake language model, which returns scripted answers, so their output is real but the answers are staged. Samples that need a real model or training hardware say so in the sentence before them, and any result they show is an illustration.
- Two models. The task model is a small local model served by Ollama, run many times. The reflection model (for GEPA) and the teacher (for bootstrapping and fine-tuning) is Claude Opus 5.5, used sparingly. The pages explain why this split is the usual one. One exception: DSPy cannot fine-tune a model served by Ollama, so the weight pages in stage 4 train a Hugging Face copy of a small model instead and say so.
- Scores always carry their split and size. "62% on a 50-example valset", never a bare "62%".
- Words with fixed meanings: a program is a DSPy module tree, a predictor is one model call inside it, a candidate is one version of the program during optimization, a rollout is one run on one example, and a trace is the record of every predictor's inputs and outputs during a rollout.
- Hardware is described generically, as "a GPU with N GB of memory", with numbers from the published documentation.
- There is no shared running example yet; each page uses its own small task.
Recommended software
Names only, no links: names stay stable, download pages do not.
| Job | Pick | Why |
|---|---|---|
| The framework | DSPy | Signatures, modules, metrics, evaluation, and every optimizer in this cluster. |
| GEPA outside DSPy | gepa (the standalone package) | The same optimizer for any text you can score: a system prompt, a config, a piece of code. DSPy installs it for you. |
| Running the local task model | Ollama | Serves open models on your own machine behind a simple API that DSPy can call. |
| Seeing what the program did | MLflow tracing, or DSPy's own inspect_history |
Every prompt sent and every answer received, which is how you debug a metric or a module. |
| Writing and running experiments | Any Python editor with a terminal, or a notebook | Optimization runs are long; a notebook keeps results around between steps. |
Prices and versions are deliberately absent, apart from the DSPy version the samples were checked against.
How to read this cluster
The learning path below is sorted so that each row depends only on the rows above it. If you already know how language models work and have written Python, start at Signatures and modules. If you came for GEPA, start at its algorithm page and follow its links back only as far as you need.
Learning path
Each row depends only on rows above it. Read top to bottom, or jump to a technique and follow its "Before this" links downward.
- prerequisiteSearch and optimization basicsAn objective, a space of candidates, and a strategy for trying them: random search, hill climbing, and the trade between exploring and exploiting.
- prerequisiteHow language models generate textTokens, the context window, next-token probabilities, and how temperature and sampling turn one prompt into different outputs.
- prerequisiteTraining, validation, and test splitsWhy you split examples into sets with different jobs, and how using the wrong set gives a score that lies.
- prerequisitePython essentials for DSPyJust enough Python to read and run every sample in this cluster: classes and inheritance, type hints, keyword arguments, callables, and virtual environments.
- prerequisiteEvolutionary algorithmsKeeping a population of candidates, choosing parents, making mutated children, and keeping the good ones: the family of search methods GEPA belongs to.
- prerequisiteGradient descent and fine-tuningHow a model's weights change during training: a loss, its slope, a learning rate, and repeated small steps, and what fine-tuning changes compared with prompting.
- prerequisiteNoisy scores and sample sizeWhy a score measured on a small set wobbles, how to estimate the wobble, and how many examples you need before a difference is real.
- prerequisitePareto frontsHow to compare candidates that are each best at different things: domination, the Pareto set, and why keeping it preserves useful variety.
- prerequisitePrompts and few-shot learningWhat a prompt is made of: instructions, input fields, and worked examples, and why a few good examples can change a model's answers more than more instructions.
- prerequisiteStructured output and JSONWhy programs need model output in fields instead of free text, what JSON looks like, and what happens when a model's output does not parse.
- techniqueConfiguring language modelsConnecting DSPy to a local task model served by Ollama and to Claude Opus 5.5 for reflection, with settings, contexts, caching, and usage tracking.
- prerequisiteLoRA and parameter-efficient tuningFine-tuning a large model by training a small add-on instead of every weight: low-rank adapters, their size, and what they cost in memory.
- prerequisiteReinforcement learning basicsLearning from a reward instead of from correct answers: policies, rewards, advantages, and why comparing samples to each other gives a learning signal.
- techniqueSignatures and modulesWriting a task as typed inputs and outputs instead of a prompt string, and running it through Predict, ChainOfThought, and other built-in modules.
- techniqueAdapters and structured outputHow DSPy turns a signature into the actual messages a model sees and parses the reply back into typed fields: ChatAdapter, JSONAdapter, and what to do when parsing fails.
- techniqueBuilding a datasetTurning examples into dspy.Example objects, marking which fields are inputs, deciding how many you need, and splitting them for optimization.
- techniqueComposing programsBuilding a multi-step program as a dspy.Module subclass: several predictors wired together in forward, with tools, loops, and branches in plain Python.
- techniqueEvaluating programsRunning a program over a dataset with dspy.Evaluate, reading the results, and deciding whether a change is a real improvement or noise.
- techniqueGEPA: reflective prompt evolutionHow GEPA improves prompts by reading what went wrong: run a minibatch, reflect on traces and feedback in plain language, propose a new instruction, and keep a Pareto front of candidates.
- techniqueLLM-as-judge metricsWhen there is no single right answer: using a model with a rubric as the metric, writing the judge as a DSPy program, and keeping it honest.
- techniqueFew-shot bootstrappingThe first family of optimizers: LabeledFewShot, BootstrapFewShot, and random search over bootstrapped demos, which pick worked examples for every predictor automatically.
- techniqueGEPA in practiceRunning dspy.GEPA: the metric with feedback, the reflection model, budgets, multi-module programs, reading the results, and the standalone gepa library for optimizing any text.
- techniqueBootstrapFinetuneDistilling a program into a smaller model's weights: collecting traces that pass the metric with a strong teacher, then fine-tuning the local task model on them.
- techniqueMIPROv2Optimizing instructions and examples together: proposing candidate instructions, bootstrapping demo sets, and searching combinations with Bayesian optimization.
- techniqueSaving and serving programsKeeping an optimized program: saving state versus the whole program, loading it in production, pinning models and versions, and serving it behind an API.
- techniqueOverfitting and re-optimizingKnowing when an optimized program has stopped being good: overfitting to the valset, the model changing underneath you, data drift, and when to run the optimizer again.
- techniqueReinforcement learning with GRPOTraining the task model's weights directly against your metric: group-relative advantages, DSPy's GRPO optimizer, and why it costs far more rollouts than GEPA.