Glossary · Evaluation & safety

LLM-as-a-Judge

Using a language model to score, compare, classify, or critique another system's output against a rubric.

Why it matters

It can scale evaluation of qualities that are difficult to express as exact-match tests, such as clarity or instruction adherence.

In practice

Give a separate evaluator model the task, candidate answer, reference evidence, and a structured rubric, then calibrate its scores against human-reviewed examples.

Common confusion

A judge model is not ground truth. It can be biased by order, verbosity, style, prompt wording, or shared model failures.

Related terms

Browse the learning paths to see this term in context — every lesson is free to read.