Cookbook · RLHF
RLHF (Reinforcement Learning from Human Feedback)
4 minrlhfpporeward-modelreinforcement-learning
The idea, in one analogy
Imagine training someone by first having judges rate pairs of their attempts against each other, distilling those ratings into a general sense of "what good looks like," and then having the person practice repeatedly against that internalized judge, getting nudged whenever they stray too far from how they used to behave. RLHF does exactly this in two stages: train a reward model to imitate human preference judgments, then use reinforcement learning to optimize the language model against that learned reward, while a penalty keeps it from drifting too far from where it started.
The two stages
Stage 1: train a reward model. Collect human comparisons: given a prompt and two model-generated responses, a human (or, increasingly, another AI model) picks the better one. Train a separate model to predict that preference, so it can score any response with a single number representing "how much would a human like this."
Stage 2: optimize the policy against that reward, with PPO. Generate responses from the current policy, score them with the reward model from stage 1, and update the policy (via Proximal Policy Optimization, the standard RL algorithm here) to make highly-rewarded responses more likely. Crucially, a KL-divergence penalty against the original (pre-RL) policy is added to the objective, so the model can't drift arbitrarily far in pursuit of reward, since an unconstrained policy will often find degenerate outputs that fool the reward model without being good.
Why this is expensive and finicky
RLHF needs three models in memory during training: the policy being updated, the frozen reference policy (for the KL term), and the reward model (to score generated responses), plus the reward model needed its own separate training run beforehand on human comparison data. PPO itself is sensitive to hyperparameters and prone to instability if the reward signal is noisy or the KL penalty is miscalibrated. This is the cost DPO is designed to avoid (no separate reward model, no RL rollout loop) and GRPO is designed to reduce (no separate critic model, using group-relative scoring instead).
When RLHF is still the right call
When there's no cheap, verifiable way to score a response (the kind GRPO needs) and preference data is naturally comparative rather than pointwise (humans find it easier to compare two responses than assign an absolute score, which is exactly what the reward model in stage 1 is built to learn from), RLHF's extra machinery buys something DPO's simpler pairwise loss doesn't: a reward model that generalizes to responses it never saw a direct comparison for, since it learned a general notion of quality rather than memorizing specific pairs.
Where to look further
- Christiano et al., "Deep Reinforcement Learning from Human Preferences": the foundational RLHF paper.
- Ouyang et al., "Training language models to follow instructions with human feedback" (InstructGPT): RLHF applied to instruction-following at scale, the recipe most modern chat models descend from.
- Hugging Face's TRL library: includes maintained reward-model training and PPO implementations.
- DPO and GRPO: the two main alternatives that remove different pieces of this pipeline's cost.