Skip to content
Ishaan Reddy

Cookbook · FineTuning

QLoRA

3 minfine-tuningqloraquantizationpeft

The idea, in one analogy

LoRA already shrinks the trainable part of fine-tuning down to a small adapter. QLoRA asks the next question: what about the frozen part? The base model's weights still have to sit in GPU memory even though they never change during training, so QLoRA compresses them down to 4 bits each, the way you might vacuum-pack clothes you're not using right now but still need on hand. The compressed weights take a quarter the space of a normal 16-bit copy, and only get "unpacked" briefly, on the fly, for each computation.

Why this compounds with LoRA

LoRA's memory savings come from not needing gradients or optimizer state for the frozen weights, only for the small adapter. But the frozen weights themselves still have to be stored somewhere, and at standard 16-bit precision, a 70B-parameter model's weights alone take 140GB, out of reach for a single consumer GPU regardless of how cheap the adapter is. QLoRA quantizes those frozen weights to 4-bit, cutting that storage to roughly 35GB, while keeping the small LoRA adapter (A and B in LoRA's notation) at full precision so training quality doesn't suffer. The two techniques solve different halves of the same memory problem: LoRA shrinks what you train, QLoRA shrinks what you have to store to train it.

What this buys you, concretely

Dettmers et al.'s original benchmark: fine-tuning a 65B-parameter model, which would need over 780GB of GPU memory for full 16-bit fine-tuning, dropped to fitting on a single 48GB GPU with QLoRA, while matching full 16-bit fine-tuning performance on their benchmarks. That's the headline result: 4-bit QLoRA fine-tuning reaching parity with 16-bit full fine-tuning, not just "close enough."

Where to look further

  • Dettmers et al., "QLoRA: Efficient Finetuning of Quantized LLMs": the original paper, including the NF4 datatype and double quantization.
  • bitsandbytes: the library implementing NF4 quantization and paged optimizers, used by Hugging Face's PEFT/Transformers integration.
  • LoRA: the adapter mechanism QLoRA builds on top of.