Cookbook · Evaluation
Human evaluation
3 minevaluationhuman-evaluationinter-rater-agreement
The idea, in one analogy
Automated benchmarks (see benchmarks) and LLM judges (see LLM-as-judge) are both proxies, cheaper and faster stand-ins for the thing you care about: whether a real person finds the model's output good. Human evaluation is going back to that source directly, at the cost of the speed and scale the proxies bought you.
When the proxies aren't enough
Multiple-choice benchmarks can't measure open-ended quality (tone, helpfulness, whether an explanation is clear to a non-expert). LLM judges approximate human judgment but inherit measurable biases (see LLM-as-judge). For a launch decision, a safety-sensitive claim, or any result you plan to make a real decision on, at least a sample of direct human review is worth the cost, even if the bulk of iteration happens on cheaper automated signals.
Common human-eval setups
Pairwise preference: show a rater two responses to the same prompt (from two different models, or two versions of the same model) and ask which is better, sometimes with an option for "tie." This mirrors preference dataset formats' comparative format, just with a human rater instead of an AI judge or automated collection.
Rubric-based rating: raters score a single response against explicit criteria (accuracy, helpfulness, safety, tone), often on a numeric or Likert scale, producing a more granular signal than a single win/loss but requiring clearer rater instructions to stay consistent across raters.
A/B testing in production: rather than asking raters to evaluate outputs directly, expose different model versions to real users and measure downstream behavior (task completion, follow-up questions, explicit feedback like thumbs up/down). This measures actual user outcomes rather than a rater's judgment of quality, at the cost of needing real traffic and a longer feedback loop.
The part that's easy to get wrong: inter-rater agreement
A single rater's opinion is a single opinion, not a fact. Using multiple raters per item and measuring how often they agree (a common metric: Cohen's kappa, which corrects for the agreement you'd expect from chance alone) tells you whether the rating task itself is well-specified. Low agreement usually means the rubric is ambiguous or the task is inherently subjective, either of which should change how much weight the resulting scores get, not just be averaged away as noise.
Where to look further
- LLM-as-judge: the cheaper, faster proxy this page is the ground-truth check against.
- Evaluation: the automated, fixed-answer alternative for tasks where one exists.