Skip to content
Ishaan Reddy

Cookbook · RLHF

GRPO (Group Relative Policy Optimization)

3 mingrporeinforcement-learningrlhfreasoning-models

The idea, in one analogy

Standard reinforcement learning (PPO, the method behind classic RLHF) usually needs a "critic": a second model trained alongside the policy to estimate how good a given state is, so the policy knows whether a reward was better or worse than expected. That's like grading a single essay against your own private sense of what a good essay looks like, a sense you have to keep separately trained and calibrated. GRPO skips the critic entirely: for a given prompt, it generates a whole group of candidate responses, scores all of them, and simply compares each one to the average of its own group. No separate model needed to say what "average" should have been; the group defines it.

Why this matters

PPO's critic model roughly doubles the memory and compute cost of RL training (a second full model, trained simultaneously, whose own errors can destabilize the policy update). For tasks where you can cheaply generate several candidate responses per prompt and score them (math problems with a checkable answer, code that either passes tests or doesn't), GRPO trades that critic model for extra sampling: generate a group of, say, 8-16 responses per prompt, score each one, and use each response's score relative to its group's mean as the training signal. This is the RL method behind DeepSeek-R1-style reasoning training, where "does this math answer match the known correct one" gives a cheap, reliable per-response reward without needing any learned reward or value model at all.

What this needs that DPO and RLHF don't

GRPO needs to sample multiple responses per prompt during training (not just one), and it needs some way to score each one, either a verifiable checker (does this code pass the unit tests, does this equation evaluate correctly) or a reward model. It's a strong fit when a verifiable checker exists, since it turns "was this response good" into a cheap, objective fact rather than something a reward model has to learn to approximate. It's a weaker fit for open-ended tasks (creative writing, general helpfulness) where there's no cheap way to check correctness and you'd need a reward model anyway, at which point RLHF or DPO become more natural starting points.

Where to look further

  • Shao et al., "DeepSeekMath": the paper that introduced GRPO.
  • Guo et al., "DeepSeek-R1": GRPO applied at scale to reasoning-focused RL training.
  • RLHF: the critic-based alternative GRPO is designed to avoid the cost of.