Cookbook · RLHF
Preference dataset formats
3 minpreference-datadporlhfdatasets
The idea, in one analogy
An SFT dataset (see SFT dataset formats) tells a model "here is the one right answer." A preference dataset instead says "between these two answers, this one was better," without necessarily claiming either is perfect. That comparative framing is what DPO and RLHF are built to learn from.
The basic shape: a triple
A preference example is a prompt with two candidate responses, one marked as preferred:
{"prompt": "Explain photosynthesis to a 10-year-old.",
"chosen": "Plants use sunlight to turn water and air into food and oxygen, like a tiny food factory powered by the sun.",
"rejected": "Photosynthesis is the process by which photoautotrophs convert light energy into chemical energy via chlorophyll-mediated reactions."}This is exactly the (y_w, y_l) pair in DPO's loss, and exactly the kind of comparison RLHF's reward model is trained on.
Where these comparisons come from
Human annotation: two model outputs are shown to a human rater who picks the better one, sometimes with a rating scale or explanation attached. High quality, but slow and expensive to collect at scale.
AI-generated preferences ("RLAIF"): a stronger model judges pairs of outputs from a weaker model being trained, following a written rubric. Much cheaper to scale than human annotation, at the cost of inheriting whatever biases or blind spots the judging model has (see LLM-as-judge for the same caveat applied to evaluation instead of training data).
Rejection sampling: generate several candidate responses to the same prompt from the model itself, score them (by a reward model, a rubric, or a verifiable checker), and pair the best-scoring and worst-scoring as chosen/rejected. This is also the same sampling setup GRPO uses, just consumed differently: GRPO uses the whole group's relative scores directly, while this produces discrete pairs for DPO-style training.
A trap worth checking for
If chosen responses are systematically longer, more formatted, or more confident-sounding than rejected ones, independent of whether they're more correct or helpful, a model trained on this data learns to produce longer, more confident-sounding text rather than better text. This "length bias" is a documented, measurable failure mode in preference datasets; checking whether chosen/rejected length distributions differ substantially is a cheap sanity check before training on a preference dataset at all.
Where to look further
- DPO and RLHF: what consumes this data.
- Stiennon et al., "Learning to summarize from human feedback": one of the earlier, clearly documented preference-data collection pipelines.