A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and review procedures.
Why it matters
You cannot improve reliability if success is only a subjective impression from a few demos.
In practice
Run the same customer-support scenarios before and after changing retrieval, score correctness and citation support, and inspect failures by category.
Common confusion
A benchmark score is one evaluation result, not a complete account of production quality.
Related terms
Browse the learning paths to see this term in context — every lesson is free to read.