prerequisite

Prompts and few-shot learning

What a prompt is made of: instructions, input fields, and worked examples, and why a few good examples can change a model's answers more than more instructions.

Before this

This page assumes you are comfortable with:

Why you need this

Every prompt optimizer in this cluster changes exactly two things: the instruction and the examples in the prompt. Before you can follow what an optimizer is doing, you need to see those parts in an ordinary hand-written prompt and know what each one does to the model's answer. This page takes one small task apart, with no DSPy code.

The idea

A prompt is the text you send a language model so that its continuation is the answer you want. A language model is a program that continues text one token (a chunk of text, often a word or part of a word) at a time, picking each next token from a probability distribution that depends on everything before it (How language models generate text). So everything in the prompt shifts those probabilities, whether you meant it to or not.

A useful prompt for a well-defined task usually has four parts:

Part What it is What it does to the answer
Instruction One or more sentences describing the task Sets which kind of continuation is likely
Output format What the answer should look like: one word from a list, a number, a JSON object Makes the answer easy for code to read
Examples Worked input and output pairs, also called demonstrations Shows the pattern instead of describing it
Input The actual case to answer The part that changes on every call

Zero-shot means the prompt has no examples: instruction, format, input. Few-shot means it includes a few examples, typically between one and a handful, before the real input. The model was never retrained on them. It imitates them, because a continuation that follows the same pattern as the text above it is likely. This is called in-context learning: learning from the context window during one call, forgotten as soon as the call ends.

Chain of thought is asking the model to write its reasoning before the answer. Because the answer tokens come after the reasoning tokens, the answer is generated from a context that already contains the worked-out steps. It tends to help on tasks with several steps, such as arithmetic word problems, and it costs extra output tokens on every call.

Worked example

The task: label a customer message as question, complaint, or praise.

Zero-shot

Classify the customer message as question, complaint, or praise.
Answer with one word.

Message: The box was crushed and the lid is cracked. Do you ship replacements?
Label:

Annotated:

Line Part Note
Classify the customer message... Instruction Names the three labels, but not where the lines fall between them
Answer with one word. Output format Asks for something a program can compare with ==
Message: ... Input Contains a complaint and a question at once
Label: Cue Ends the prompt where the answer should start, so the next tokens are likely to be the label

This message is hard on purpose. It reports damage and asks a question. With nothing but the instruction, the model has to guess which label wins, and two models (or two runs with sampling) can reasonably disagree. The instruction did not say.

Few-shot

Classify the customer message as question, complaint, or praise.
Answer with one word.

Message: How long does delivery to Canada take?
Label: question

Message: My order arrived two weeks late and nobody answered my emails.
Label: complaint

Message: The parcel came damaged. Can I get a refund?
Label: complaint

Message: The box was crushed and the lid is cracked. Do you ship replacements?
Label:

What each added part does:

Added part Effect
Example 1, a plain question Anchors question to messages that only ask for information
Example 2, a plain complaint Anchors complaint to reports of something going wrong
Example 3, damage plus a question, labeled complaint Settles the hard case: a problem report wins even when it ends in a question
The repeated Message: / Label: layout Shows the exact output shape: one lowercase word, no punctuation, no explanation

The third example carries most of the weight. It encodes a decision rule ("a problem report wins over a question") that you would otherwise need a sentence of instruction to state, and that sentence might still be read differently by a different model. That is the sense in which a few good examples can change answers more than more instructions: they show the boundary between labels at the place it is hardest to describe.

Examples also have costs and side effects. If all three had been complaint, the model would lean toward complaint for everything. If the examples were very long, every call would pay for those tokens. Choosing which examples to show is a search problem, and it is one of the two things DSPy's optimizers automate.

With chain of thought

To ask for reasoning first, change the output format and the cue:

Classify the customer message as question, complaint, or praise.
First write one sentence of reasoning, then the label.

Message: The box was crushed and the lid is cracked. Do you ship replacements?
Reasoning:

A model might now write "The customer reports damage and asks for a fix, which is a complaint." before Label: complaint. Your code must find the label after the reasoning, which is a parsing job (Structured output and JSON).

Brittleness, which motivates DSPy

A hand-written prompt is tuned to one model. Change any of these and the tuning can break:

  • The model. A new or smaller model reads the same instruction differently, or ignores the one-word format and adds a sentence.
  • The task. A fourth label is added, and the examples no longer cover every boundary.
  • The data. New kinds of messages show up that the examples never touched.

Each fix is another edit to a long string, checked by eye on a few cases. Nobody can say whether the edit helped on average, and an edit that fixes one case often breaks another. DSPy's answer is to stop treating the prompt as the thing you write. You write the task's inputs and outputs, a metric, and a set of examples; the library builds the prompt, and an optimizer chooses the instruction and the examples by measuring them.

In an optimization pipeline

In stage 1, Define the program, the instruction and output format become a signature, and the prompt is generated from it (Signatures and modules). In stage 3, Optimize the prompts, few-shot bootstrapping picks examples automatically, MIPROv2 searches instructions and examples together, and GEPA rewrites instructions after reading what went wrong. Everything on this page is what those optimizers are changing.

Common mistakes

  • Examples that all share one label. The model answers that label for almost everything.
  • Examples that disagree with the instruction. The instruction says "one word" but an example label has a period; replies start arriving with periods and exact-match checks fail.
  • Only easy examples. The examples repeat what the instruction already says, and the hard cases are still guessed.
  • An example that is also a test case. The model copies the answer, and the score looks better than it will be on new inputs.
  • Asking for reasoning without changing the parser. The reasoning sentence lands where the code expected a single word, and every answer counts as wrong.

Cost

Every token in the prompt is paid for on every call. If the instruction and format take tIt_I tokens, each example takes tEt_E tokens, there are kk examples, and the input takes txt_x tokens, the prompt is tI+k⋅tE+txt_I + k \cdot t_E + t_x tokens long, so examples add k⋅tEk \cdot t_E tokens to every call. Chain of thought adds output tokens, which many providers price higher than input tokens and which also take time to generate. Writing and checking prompts by hand costs the builder's time, and that cost repeats every time the model changes.

Going further

  • Signatures and modules, where this page's prompt becomes a DSPy program.
  • The original few-shot results in the GPT-3 paper, "Language Models are Few-Shot Learners" (Brown and colleagues, 2020).
  • The chain-of-thought prompting paper by Wei and colleagues (2022).
  • Few-shot bootstrapping, for how DSPy chooses examples for you.

Leads to

Back to DSPy and GEPA: programming and optimizing language-model systems