Cookbook · FineTuning
LoRA (Low-Rank Adaptation)
4 minfine-tuninglorapefttraining
The idea, in one analogy
Imagine editing a massive, finely-tuned machine instead of rebuilding it. Rather than rewiring every one of its million internal connections, you bolt on a small control panel with a handful of dials, and only those dials get adjusted. The machine's original wiring stays untouched; the dials just nudge its behavior. LoRA (Low-Rank Adaptation) does this for a pretrained model's weights: instead of updating a full weight matrix during fine-tuning, it freezes the matrix entirely and trains a small "adapter" bolted alongside it.
Why this matters
Full fine-tuning (see SFT) updates every parameter in a model, which for a large model means holding gradients and optimizer state for the whole thing, memory that scales with model size regardless of how narrow the actual behavior change is. LoRA freezes the pretrained weights entirely and trains a small number of added parameters instead. The base model's knowledge stays untouched; only a small, cheap-to-train adapter changes. This is what makes fine-tuning a 7B+ parameter model feasible on a single consumer GPU, since the memory-hungry part (gradients and optimizer state) only has to cover the tiny adapter, not the full model.
Where the adapter goes
LoRA is applied per weight matrix, not to the model as a whole. In a transformer (see transformers), the usual targets are the attention projections (W_Q, W_K, W_V, the output projection) and sometimes the MLP's linear layers. Applying it to more matrices increases the adapter's capacity (and its parameter count) at the cost of some of the memory savings; applying it to fewer keeps things cheaper but limits how much the model's behavior can shift.
What you get at the end
A trained LoRA adapter is a small, separate set of weights, often tens of megabytes, that can be merged into the base model's weights (W' = W + BA, computed once) for zero-overhead inference, or kept separate and swapped at load time so one base model can serve many different fine-tuned "personalities" without duplicating the whole model per task.
Where to look further
- Hu et al., "LoRA: Low-Rank Adaptation of Large Language Models": the original paper.
- Hugging Face's PEFT library: the standard tooling for LoRA fine-tuning.
- Unsloth: custom kernels that make LoRA/QLoRA fine-tuning significantly faster and lower-memory on consumer GPUs.
- QLoRA: combines this with a quantized base model for even lower memory use.