prerequisite
How language models generate text
Tokens, the context window, next-token probabilities, and how temperature and sampling turn one prompt into different outputs.
Before this
Nothing beyond first-year college math. This is a starting page.
Why you need this
Every DSPy program ends in a call to a language model, and every optimizer in this cluster runs that call hundreds or thousands of times. To read a score, set a budget, or understand why a program answered differently on the second run, you need a working picture of what the model does with the text you send it. This page builds that picture from the bottom: tokens, probabilities, and sampling.
The idea
Tokens
A language model is a program that continues text. It does not read letters or whole words. It reads tokens: chunks of text from a fixed list called the vocabulary. A tokenizer cuts your text into tokens and replaces each one with its position in the vocabulary, a whole number called the token ID. Common words are usually one token; rare words split into pieces; a leading space is often part of the token.
Here is a toy tokenizer with a tiny vocabulary. It always takes the longest vocabulary entry that matches at the current position.
| Vocabulary entry | un |
believ |
able |
cat |
s |
|---|---|---|---|---|---|
| Token ID | 0 | 1 | 2 | 3 | 4 |
| Text | Tokens | Count |
|---|---|---|
unbelievable cats |
un believ able cat s |
5 |
believable |
believ able |
2 |
So the model sees unbelievable cats as the IDs 0 1 2 3 4. Real vocabularies hold tens of thousands of entries or more and each model family has its own, so one sentence is a different number of tokens on different models. Models are priced and limited by tokens, not words.
The context window
The context window is the most tokens a model can look at in one call: your prompt plus everything it writes back. Anything beyond it is cut off or rejected. Each model documents its own window size. Instructions, worked examples, and the input all share this one budget, which is why adding examples to a prompt has a cost.
One token at a time
The precise version: a model is a function from a context (the list of tokens so far) to a probability distribution over the vocabulary, meaning one number between 0 and 1 for every possible next token, with all of them adding up to 1. Internally it first produces a raw score for every token (often called a logit), then converts scores to probabilities with the softmax function:
Here is the raw score of token , is the base of natural exponents, the sum in the bottom runs over every token in the vocabulary, and is the temperature, a positive number you choose. Dividing by the sum is what makes the probabilities add up to 1.
Generation is a loop. The model picks one next token from the distribution, appends it to the context, and runs again. It stops at a special end token or at a length limit (max_tokens in DSPy).
Greedy versus sampled decoding
Decoding is the rule for picking the token. Greedy decoding always takes the most probable token, so the same context always gives the same text. Sampling draws a token at random in proportion to its probability, like spinning a wheel whose slices are sized by the probabilities. Sampling gives variety, and it is what most chat services do by default.
Temperature reshapes the distribution
Temperature divides every score before the softmax. A low stretches the gaps between scores, so the favorite gets even more likely. A high squeezes the gaps, so unlikely tokens get a real chance. As gets close to 0, sampling behaves like greedy decoding.
Worked example
The context is The cat sat on the. Suppose the model gives four candidate tokens these raw scores (the other tokens are left out to keep the arithmetic small):
| Token | mat |
floor |
sofa |
moon |
|---|---|---|---|---|
| Score | 3 | 2 | 1 | -1 |
At the scores are unchanged. Exponentiate each one: , , , . They add to 30.56. Divide each by the total: mat gets , or 65.7%.
At the scores double to 6, 4, 2, -2 before exponentiating. At they halve to 1.5, 1, 0.5, -0.5. The full results:
| Temperature | mat |
floor |
sofa |
moon |
|---|---|---|---|---|
| 0.5 | 86.7% | 11.7% | 1.6% | 0.03% |
| 1 | 65.7% | 24.2% | 8.9% | 1.2% |
| 2 | 47.4% | 28.7% | 17.4% | 6.4% |
Read down the moon column: an absurd answer goes from almost impossible at 0.5 to about one draw in sixteen at 2.
This script computes the same table and then samples ten times at with a seeded random generator, so the run is repeatable:
# temperature.py
import math
import random
scores = {"mat": 3.0, "floor": 2.0, "sofa": 1.0, "moon": -1.0}
def softmax(scores, temperature):
scaled = {tok: s / temperature for tok, s in scores.items()}
total = sum(math.exp(v) for v in scaled.values())
return {tok: math.exp(v) / total for tok, v in scaled.items()}
for t in (0.5, 1.0, 2.0):
probs = softmax(scores, t)
print(f"T={t}: " + " ".join(f"{tok} {p:.1%}" for tok, p in probs.items()))
rng = random.Random(4)
probs = softmax(scores, 1.0)
draws = rng.choices(list(probs), weights=list(probs.values()), k=10)
print("greedy:", max(probs, key=probs.get))
print("10 samples at T=1:", " ".join(draws))
python temperature.py prints:
T=0.5: mat 86.7% floor 11.7% sofa 1.6% moon 0.0%
T=1.0: mat 65.7% floor 24.2% sofa 8.9% moon 1.2%
T=2.0: mat 47.4% floor 28.7% sofa 17.4% moon 6.4%
greedy: mat
10 samples at T=1: mat mat mat mat mat mat sofa floor floor mat
Why the same prompt gives different answers
Greedy decoding would print mat every time. Sampling at printed mat six times, floor twice, and sofa once. Now remember that a real answer is dozens of tokens, each drawn this way, and each draw changes the context for the next one. One early floor instead of mat sends the rest of the sentence somewhere new. Two runs both starting with mat happens only of the time, and two runs picking the same first token of any kind (add the square of each probability) only about 50% of the time.
Chat models are still text
A chat model takes a list of messages, each with a role: system (standing instructions), user (the person), assistant (the model's earlier replies). Underneath, the provider joins them into one token sequence with special marker tokens between turns, then asks the model to continue after an opening assistant marker. A generic illustration (every model family uses its own markers):
<|system|>You label product reviews.<|end|>
<|user|>Battery lasts two days. Love it.<|end|>
<|assistant|>
The model then generates tokens until it produces the end marker. Roles are a convention the model learned in training, not a separate channel.
In an optimization pipeline
In stage 1, Define the program, you choose the models and their settings, including temperature and max_tokens (Configuring language models). In stage 2, Measure it, sampling is a source of noise: the same program scored twice can differ (Noisy scores and sample size). In stages 3 and 4, the optimizers change only what goes into the context (instructions and examples) or the weights that produce the scores; generation itself works exactly as above. What to put in the context is the subject of Prompts and few-shot learning.
Common mistakes
- Counting words instead of tokens. A prompt that "fits" by word count overflows the context window, and the call fails or the start of the prompt is silently dropped.
- Expecting identical output at temperature 1. Two runs of the same evaluation give different scores, and a "fix" appears to work only because of a lucky draw.
- Turning temperature up for creativity on a labeling task. Labels start coming back with rare, invalid values, the
mooncolumn above. - Setting
max_tokenstoo low. The answer stops mid-sentence, and a later step that expects a complete field fails to parse it. - Treating the system message as a hard rule. It is just more tokens; a long or forceful user message can still steer the reply.
Cost
You pay per token, both ways. A call with prompt tokens and generated tokens costs dollars, where and are the provider's prices per token for input and output. A local model has no per-token price but costs time: generation runs one token per step, so a reply of tokens takes steps, and each step gets slower as the context grows. Every extra example in a prompt is paid for again on every call.
Going further
- Byte-pair encoding, the most common way real tokenizer vocabularies are built.
- Top-k and top-p (nucleus) sampling, which cut off the unlikely tail before sampling.
- Your model provider's documentation on context window size and chat templates.
Leads to
- techniqueConfiguring language modelsConnecting DSPy to a local task model served by Ollama and to Claude Opus 5.5 for reflection, with settings, contexts, caching, and usage tracking.
- prerequisiteGradient descent and fine-tuningHow a model's weights change during training: a loss, its slope, a learning rate, and repeated small steps, and what fine-tuning changes compared with prompting.
- prerequisitePrompts and few-shot learningWhat a prompt is made of: instructions, input fields, and worked examples, and why a few good examples can change a model's answers more than more instructions.
- prerequisiteStructured output and JSONWhy programs need model output in fields instead of free text, what JSON looks like, and what happens when a model's output does not parse.
Back to DSPy and GEPA: programming and optimizing language-model systems