Cookbook · FineTuning
SFT dataset formats
2 minfine-tuningsftdatasetschat-templates
The idea, in one analogy
A pretraining corpus is like a library: shelf after shelf of unstructured text, read cover to cover. An SFT dataset is more like a stack of graded worksheets: each entry is a specific question paired with the specific answer you want the model to learn to give. The format of that pairing is what this page covers.
The basic shape: instruction and response
At minimum, an SFT example (see SFT) is a pair: an instruction (or prompt) and the response the model should learn to produce for it.
{"instruction": "Summarize this paragraph in one sentence.", "input": "...", "output": "..."}Some formats split "instruction" and "input" (the task description vs. the specific content to act on); others merge them into a single prompt field. Either works; what matters is consistency across the whole dataset, since the model learns the pattern of instruction-then-response, not just individual examples.
Multi-turn conversation format
Modern chat models are trained on full conversations, not single-turn pairs, using a "chat template" that wraps each turn with role markers:
{"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What's the capital of France?"},
{"role": "assistant", "content": "Paris."}
]}During training, the loss is typically computed only on the assistant's turns, not the user's or system's, since the model isn't being trained to predict what a user would say next, only how it should respond.
Where the loss is (and isn't) computed
A common mistake is computing cross-entropy loss (see pretraining) over the entire sequence, instruction included. This trains the model to also predict the instruction text itself, which wastes training signal on tokens the model will always be given at inference time anyway, and can subtly bias the model toward regenerating instruction-like text. Masking the loss so it only applies to response tokens (setting the label to an ignored value for every non-response token) is the standard fix.
Where to look further
- Hugging Face's chat templates documentation: how the role-based conversation format gets serialized into a single token sequence per model family.
- SFT: what this data trains.
- preference dataset formats: the comparative format DPO/RLHF need instead of single (instruction, response) pairs.