prerequisite

LoRA and parameter-efficient tuning

Fine-tuning a large model by training a small add-on instead of every weight: low-rank adapters, their size, and what they cost in memory.

Before this

This page assumes you are comfortable with:

Why you need this

Stage 4 of the pipeline fine-tunes the task model on traces your program produced. Training every weight of a billion-parameter model needs a lot of GPU memory and leaves you with a full copy of the model per task. Low-rank adaptation (LoRA) trains a small add-on instead, which is why fine-tuning a local model is practical at all on a single GPU.

The idea

Why full fine-tuning is expensive

Gradient descent updates each weight (one number inside the model) using its gradient (its slope). During training, for every weight you hold:

Number Why
The weight The model itself
Its gradient Computed on each batch
Two optimizer numbers The Adam optimizer keeps a running average of each slope and of its square

That is four numbers per weight, before counting the activations the model produces along the way. A model with 7 billion weights stored at 2 bytes each (16-bit numbers) is 7×109×2=147 \times 10^9 \times 2 = 14 GB for the weights alone, here with 1 GB = 1,000,000,000 bytes, and full fine-tuning needs several times that.

A matrix is a grid of numbers

Most of a language model's weights sit in matrices: rectangular grids of numbers. A d×dd \times d matrix WW has dd rows and dd columns, so d2d^2 numbers. When the model processes a token, it multiplies a list of dd numbers by WW to get a new list of dd numbers. Fine-tuning changes WW into W+ΔWW + \Delta W, where ΔW\Delta W (delta W) is the change, also d×dd \times d.

The low-rank trick

LoRA's bet, from the LoRA paper by Hu and colleagues (2021), is that the change needed for one task is simple, so ΔW\Delta W can be written as the product of two thin matrices:

ΔW=BA\Delta W = B A

where BB is d×rd \times r (tall and thin), AA is r×dr \times d (short and wide), and the rank rLoRAr_{\text{LoRA}} (written rr in formulas below) is a small number you choose, such as 8 or 16. Multiplying them gives a full d×dd \times d grid, but it is built from only 2dr2 d r numbers. During training, WW is frozen (never changed) and only AA and BB get gradients. BB starts at all zeros, so at step 0 the product is zero and the model behaves exactly like the original.

With tiny numbers, d=4d = 4 and r=1r = 1:

B=(120−1),A=(0.5010),BA=(0.501010200000−0.50−10)B = \begin{pmatrix}1\\2\\0\\-1\end{pmatrix}, \quad A = \begin{pmatrix}0.5 & 0 & 1 & 0\end{pmatrix}, \quad BA = \begin{pmatrix}0.5 & 0 & 1 & 0\\ 1 & 0 & 2 & 0\\ 0 & 0 & 0 & 0\\ -0.5 & 0 & -1 & 0\end{pmatrix}

Each entry of BABA is a row number from BB times a column number from AA. The product has 16 entries, but only 8 numbers produced it. Every row of BABA is a multiple of the same row AA; that is what "rank 1" means. Rank 8 allows 8 independent patterns.

Worked example

Parameter counts for one square matrix, full fine-tuning against LoRA. The trainable count for LoRA is 2dr2dr.

Matrix size dd Full: d2d^2 LoRA r=8r = 8: 2⋅d⋅82 \cdot d \cdot 8 LoRA r=16r = 16: 2⋅d⋅162 \cdot d \cdot 16
4 16 not possible: the rank cannot exceed dd (the r=1r = 1 example above uses 8) not possible
4096 16,777,216 65,536 (0.39%) 131,072 (0.78%)

At d=4096d = 4096, step by step: 40962=16,777,2164096^2 = 16{,}777{,}216; 2×4096×8=65,5362 \times 4096 \times 8 = 65{,}536, which is 65,536/16,777,216≈0.39%65{,}536 / 16{,}777{,}216 \approx 0.39\%; and doubling the rank to 16 doubles the count to 131,072, about 0.78%. The rank grows the adapter in a straight line, while the full matrix grows with the square of dd.

Now a whole model, as an illustration of the arithmetic only (no real model has exactly this shape on this page): suppose LoRA is applied to two 4096×40964096 \times 4096 matrices in each of 32 layers, so 64 matrices. Full fine-tuning of those matrices trains 64×16,777,216=1,073,741,82464 \times 16{,}777{,}216 = 1{,}073{,}741{,}824 weights. LoRA at r=16r = 16 trains 64×131,072=8,388,60864 \times 131{,}072 = 8{,}388{,}608. Stored at 2 bytes each, the adapter is 8,388,608×2≈16.88{,}388{,}608 \times 2 \approx 16.8 MB, small enough to keep one per task.

The published results for real models: the LoRA paper reports that, compared with fully fine-tuning GPT-3 175B with Adam, LoRA reduces trainable parameters by 10,000 times and GPU memory needed by 3 times.

Adapters as swappable files

Because WW never changes, the trained AA and BB for every adapted matrix can be saved on their own as an adapter. One base model on disk can serve many tasks: load the base once, then attach the ticket-triage adapter or the summarizer adapter. For speed you can also merge an adapter into the base by computing W+BAW + BA once and saving the result, which removes the extra multiplication at run time but gives up swapping.

Quantized base models

Quantization stores each frozen weight in fewer bits, for example 4 bits (half a byte) instead of 16. The 7-billion-weight model above shrinks from 14 GB to 7×109×0.5=3.57 \times 10^9 \times 0.5 = 3.5 GB of weights. Since LoRA never updates the base, the base can stay quantized while the small AA and BB train at full precision. This combination is QLoRA (Dettmers and colleagues, 2023), whose paper reports fine-tuning a 65-billion-parameter model on a single GPU with 48 GB of memory while preserving full 16-bit fine-tuning performance on its tasks.

In an optimization pipeline

In stage 4, BootstrapFinetune turns metric-filtered traces into training pairs and fine-tunes a local model. DSPy 3.4's local training provider trains a checkpoint with the PyTorch and Hugging Face libraries, and its use_peft option switches it to LoRA with rank 32. The model trained is one whose weights you hold as files, not one served by Ollama. With LoRA, the result you deploy is the base plus an adapter. Reinforcement learning with GRPO also changes weights, and the same memory arithmetic applies to it.

Common mistakes

  • Rank too low for the task. A rank of 1 or 2 may not have room for the change you need. Symptom: the tuned model scores no better than the base on your valset.
  • Rank too high on little data. More trainable numbers on a few hundred examples overfit faster. Symptom: training loss near zero, valset score dropping.
  • Loading the adapter on the wrong base. An adapter only fits the exact base model it was trained on. Symptom: a shape error on load, or fluent nonsense if the shapes happen to match.
  • Forgetting the quantization cost. A 4-bit base is smaller but slightly less accurate. Measure the quantized base before tuning, so you know which change moved the score.

Cost

Let PP be the number of weights in the base model, bb the bytes per stored weight, mm the number of adapted d×dd \times d matrices, and rr the rank. Memory for the frozen base is about P⋅bP \cdot b (14 GB at P=7×109P = 7 \times 10^9, b=2b = 2; 3.5 GB at b=0.5b = 0.5). Trainable numbers are 2drm2 d r m, and only those carry gradients and optimizer state, so the training overhead is a few times 2drm2 d r m instead of a few times PP. The adapter file is 2drm⋅b2 d r m \cdot b bytes. Training time still scales with the number of training tokens, because every token must pass through the full base model. Hardware stays generic here: you need a GPU with enough memory for the base at your chosen precision plus activations, and DSPy 3.4's local trainer uses an NVIDIA GPU (CUDA) if it finds one, then an Apple GPU, then the CPU, which is far slower.

Going further

  • The LoRA paper, "LoRA: Low-Rank Adaptation of Large Language Models" (Hu and colleagues, 2021).
  • The QLoRA paper, "QLoRA: Efficient Finetuning of Quantized LLMs" (Dettmers and colleagues, 2023).
  • BootstrapFinetune, for where DSPy uses all of this.
  • The idea of matrix rank in any first linear algebra course, which explains why a product of thin matrices can only express a few independent patterns.

Leads to

Back to DSPy and GEPA: programming and optimizing language-model systems