Skip to content
Ishaan Reddy

Cookbook · RLHF

RLHF (Reinforcement Learning from Human Feedback)

4 minrlhfpporeward-modelreinforcement-learning

The idea, in one analogy

Imagine training someone by first having judges rate pairs of their attempts against each other, distilling those ratings into a general sense of "what good looks like," and then having the person practice repeatedly against that internalized judge, getting nudged whenever they stray too far from how they used to behave. RLHF does exactly this in two stages: train a reward model to imitate human preference judgments, then use reinforcement learning to optimize the language model against that learned reward, while a penalty keeps it from drifting too far from where it started.

The two stages

Stage 1: train a reward model. Collect human comparisons: given a prompt and two model-generated responses, a human (or, increasingly, another AI model) picks the better one. Train a separate model to predict that preference, so it can score any response with a single number representing "how much would a human like this."

Stage 2: optimize the policy against that reward, with PPO. Generate responses from the current policy, score them with the reward model from stage 1, and update the policy (via Proximal Policy Optimization, the standard RL algorithm here) to make highly-rewarded responses more likely. Crucially, a KL-divergence penalty against the original (pre-RL) policy is added to the objective, so the model can't drift arbitrarily far in pursuit of reward, since an unconstrained policy will often find degenerate outputs that fool the reward model without being good.

Why this is expensive and finicky

RLHF needs three models in memory during training: the policy being updated, the frozen reference policy (for the KL term), and the reward model (to score generated responses), plus the reward model needed its own separate training run beforehand on human comparison data. PPO itself is sensitive to hyperparameters and prone to instability if the reward signal is noisy or the KL penalty is miscalibrated. This is the cost DPO is designed to avoid (no separate reward model, no RL rollout loop) and GRPO is designed to reduce (no separate critic model, using group-relative scoring instead).

When RLHF is still the right call

When there's no cheap, verifiable way to score a response (the kind GRPO needs) and preference data is naturally comparative rather than pointwise (humans find it easier to compare two responses than assign an absolute score, which is exactly what the reward model in stage 1 is built to learn from), RLHF's extra machinery buys something DPO's simpler pairwise loss doesn't: a reward model that generalizes to responses it never saw a direct comparison for, since it learned a general notion of quality rather than memorizing specific pairs.

Where to look further