Cookbook · FineTuning
Supervised fine-tuning (SFT)
3 minfine-tuningsfttrainingpeft
The idea, in one analogy
Pretraining is like getting a general education: years of broad reading to become fluent and knowledgeable in general. Fine-tuning is a focused apprenticeship afterward: a few weeks with an expert, on a narrower set of examples, to adapt what you already know to a specific job. It's fast precisely because you're not starting from zero.
Why it's cheap relative to pretraining
Pretraining (see pretraining) has to teach a model language itself, from randomly-initialized weights, which requires enormous amounts of data and compute. Fine-tuning starts from a model that already has that foundation, so it needs orders of magnitude less data (often thousands to low millions of examples, not billions of tokens) and far less compute to noticeably shift behavior toward a task, domain, or style.
"Supervised" here means the training data is (input, desired output) pairs, the same kind of labeled data used in ordinary supervised learning, just applied to a pretrained model instead of a randomly-initialized one. A typical SFT example is a (instruction, ideal response) pair: the model is trained to produce the response given the instruction, using the same next-token cross-entropy loss as pretraining, just computed only over the response tokens.
Full fine-tuning updates every parameter in the model, using the same mechanics as pretraining, just on a smaller, more targeted dataset and usually a much smaller learning rate. It's effective, but still requires enough memory to hold gradients and optimizer state for the entire model. Parameter-efficient fine-tuning (PEFT), covered in LoRA and QLoRA, is the common alternative when that memory cost is prohibitive.
When to fine-tune instead of pretraining from scratch
If the capability you need (general language understanding, broad world knowledge) already exists in an available pretrained model, and what you need is a narrower behavior change (a specific tone, a specific task format, domain vocabulary), fine-tuning gets you there for a tiny fraction of the cost of pretraining. Training from scratch pays off when the model's foundation itself must differ: a different tokenizer/vocabulary for a distinct domain or language, a novel architecture, or when the point is to learn how the whole pipeline works rather than to ship a product as fast as possible.
A failure mode to watch for
Fine-tuning on a narrow dataset for too long, or too aggressively, degrades the broad capability the base model started with: the model gets better at the narrow task and worse at everything else it used to be able to do ("catastrophic forgetting"). Mixing in some general-purpose data alongside the task-specific data during fine-tuning, and keeping the learning rate low, are the standard mitigations.
Where to look further
- Hugging Face's TRL library: the standard tooling for SFT, along with preference tuning (see DPO, RLHF).
- LoRA and QLoRA: the parameter-efficient alternative to full fine-tuning, for when memory is the constraint.