prerequisite
LoRA and parameter-efficient tuning
Fine-tuning a large model by training a small add-on instead of every weight: low-rank adapters, their size, and what they cost in memory.
Before this
This page assumes you are comfortable with:
Why you need this
Stage 4 of the pipeline fine-tunes the task model on traces your program produced. Training every weight of a billion-parameter model needs a lot of GPU memory and leaves you with a full copy of the model per task. Low-rank adaptation (LoRA) trains a small add-on instead, which is why fine-tuning a local model is practical at all on a single GPU.
The idea
Why full fine-tuning is expensive
Gradient descent updates each weight (one number inside the model) using its gradient (its slope). During training, for every weight you hold:
| Number | Why |
|---|---|
| The weight | The model itself |
| Its gradient | Computed on each batch |
| Two optimizer numbers | The Adam optimizer keeps a running average of each slope and of its square |
That is four numbers per weight, before counting the activations the model produces along the way. A model with 7 billion weights stored at 2 bytes each (16-bit numbers) is GB for the weights alone, here with 1 GB = 1,000,000,000 bytes, and full fine-tuning needs several times that.
A matrix is a grid of numbers
Most of a language model's weights sit in matrices: rectangular grids of numbers. A matrix has rows and columns, so numbers. When the model processes a token, it multiplies a list of numbers by to get a new list of numbers. Fine-tuning changes into , where (delta W) is the change, also .
The low-rank trick
LoRA's bet, from the LoRA paper by Hu and colleagues (2021), is that the change needed for one task is simple, so can be written as the product of two thin matrices:
where is (tall and thin), is (short and wide), and the rank (written in formulas below) is a small number you choose, such as 8 or 16. Multiplying them gives a full grid, but it is built from only numbers. During training, is frozen (never changed) and only and get gradients. starts at all zeros, so at step 0 the product is zero and the model behaves exactly like the original.
With tiny numbers, and :
Each entry of is a row number from times a column number from . The product has 16 entries, but only 8 numbers produced it. Every row of is a multiple of the same row ; that is what "rank 1" means. Rank 8 allows 8 independent patterns.
Worked example
Parameter counts for one square matrix, full fine-tuning against LoRA. The trainable count for LoRA is .
| Matrix size | Full: | LoRA : | LoRA : |
|---|---|---|---|
| 4 | 16 | not possible: the rank cannot exceed (the example above uses 8) | not possible |
| 4096 | 16,777,216 | 65,536 (0.39%) | 131,072 (0.78%) |
At , step by step: ; , which is ; and doubling the rank to 16 doubles the count to 131,072, about 0.78%. The rank grows the adapter in a straight line, while the full matrix grows with the square of .
Now a whole model, as an illustration of the arithmetic only (no real model has exactly this shape on this page): suppose LoRA is applied to two matrices in each of 32 layers, so 64 matrices. Full fine-tuning of those matrices trains weights. LoRA at trains . Stored at 2 bytes each, the adapter is MB, small enough to keep one per task.
The published results for real models: the LoRA paper reports that, compared with fully fine-tuning GPT-3 175B with Adam, LoRA reduces trainable parameters by 10,000 times and GPU memory needed by 3 times.
Adapters as swappable files
Because never changes, the trained and for every adapted matrix can be saved on their own as an adapter. One base model on disk can serve many tasks: load the base once, then attach the ticket-triage adapter or the summarizer adapter. For speed you can also merge an adapter into the base by computing once and saving the result, which removes the extra multiplication at run time but gives up swapping.
Quantized base models
Quantization stores each frozen weight in fewer bits, for example 4 bits (half a byte) instead of 16. The 7-billion-weight model above shrinks from 14 GB to GB of weights. Since LoRA never updates the base, the base can stay quantized while the small and train at full precision. This combination is QLoRA (Dettmers and colleagues, 2023), whose paper reports fine-tuning a 65-billion-parameter model on a single GPU with 48 GB of memory while preserving full 16-bit fine-tuning performance on its tasks.
In an optimization pipeline
In stage 4, BootstrapFinetune turns metric-filtered traces into training pairs and fine-tunes a local model. DSPy 3.4's local training provider trains a checkpoint with the PyTorch and Hugging Face libraries, and its use_peft option switches it to LoRA with rank 32. The model trained is one whose weights you hold as files, not one served by Ollama. With LoRA, the result you deploy is the base plus an adapter. Reinforcement learning with GRPO also changes weights, and the same memory arithmetic applies to it.
Common mistakes
- Rank too low for the task. A rank of 1 or 2 may not have room for the change you need. Symptom: the tuned model scores no better than the base on your valset.
- Rank too high on little data. More trainable numbers on a few hundred examples overfit faster. Symptom: training loss near zero, valset score dropping.
- Loading the adapter on the wrong base. An adapter only fits the exact base model it was trained on. Symptom: a shape error on load, or fluent nonsense if the shapes happen to match.
- Forgetting the quantization cost. A 4-bit base is smaller but slightly less accurate. Measure the quantized base before tuning, so you know which change moved the score.
Cost
Let be the number of weights in the base model, the bytes per stored weight, the number of adapted matrices, and the rank. Memory for the frozen base is about (14 GB at , ; 3.5 GB at ). Trainable numbers are , and only those carry gradients and optimizer state, so the training overhead is a few times instead of a few times . The adapter file is bytes. Training time still scales with the number of training tokens, because every token must pass through the full base model. Hardware stays generic here: you need a GPU with enough memory for the base at your chosen precision plus activations, and DSPy 3.4's local trainer uses an NVIDIA GPU (CUDA) if it finds one, then an Apple GPU, then the CPU, which is far slower.
Going further
- The LoRA paper, "LoRA: Low-Rank Adaptation of Large Language Models" (Hu and colleagues, 2021).
- The QLoRA paper, "QLoRA: Efficient Finetuning of Quantized LLMs" (Dettmers and colleagues, 2023).
- BootstrapFinetune, for where DSPy uses all of this.
- The idea of matrix rank in any first linear algebra course, which explains why a product of thin matrices can only express a few independent patterns.
Leads to
Back to DSPy and GEPA: programming and optimizing language-model systems