Skip to content
Ishaan Reddy

Cookbook · Evaluation

LLM-as-judge

3 minevaluationllm-as-judgebias

The idea, in one analogy

Grading an open-ended essay is harder than grading a multiple-choice test: there's no single correct string to check against, only a judgment call about quality. Evaluation's multiple-choice scoring sidesteps this by only asking questions with a known right answer. LLM-as-judge tackles the harder case, open-ended generation, by using a strong model to play the role of the grader, given a rubric, the way a teaching assistant might grade essays against a marking scheme instead of an answer key.

How it works

A capable model (often, though not always, a larger or more capable one than the model being evaluated) is given the prompt, the response being graded, and a rubric or set of criteria, then asked to produce a score or a verdict. Common patterns:

  • Pointwise scoring: grade a single response on a scale (1-10, or a rubric with sub-criteria), independent of any other response.
  • Pairwise comparison: given two responses to the same prompt, judge which is better, the same comparative framing preference dataset formats uses to build training data, just applied to evaluation instead.
  • Reference-guided grading: give the judge a known-good reference answer alongside the response being graded, so it has something concrete to compare against rather than judging from first principles.

Why this is useful

It scales to open-ended tasks (summarization quality, helpfulness, following complex instructions) where a fixed-answer benchmark (see Evaluation) simply doesn't apply, and it's far cheaper and faster than collecting human judgments for every evaluation run.

The caveat that matters

An LLM judge is not a ground truth, it's a proxy for human judgment, and proxies can diverge from what they're standing in for. Documented failure modes: judges tend to favor longer responses regardless of quality (the same length bias covered in preference dataset formats), can be sensitive to the order two responses are presented in (position bias), and can favor responses written in a style similar to their own outputs. Before trusting an LLM-judge pipeline's numbers, it's worth spot-checking a sample of its judgments against actual human review, ideally reporting the agreement rate rather than assuming it's high.

Where to look further