A versioned collection of inputs, expected properties, scoring rules, and metadata used to measure an AI system against a defined capability or risk.
Why it matters
A repeatable set turns vague quality claims into comparable evidence and catches regressions after prompts, models, tools, or retrieval change.
In practice
Keep representative support questions, adversarial instructions, expected citations, and failure labels in a reviewed dataset that is separate from development examples.
Common confusion
A development eval guides iteration, a final held-out test estimates performance after choices are fixed, and a standardized benchmark supports comparison under a shared protocol. Repeated tuning against any held-out set leaks test information and inflates results.
Related terms
Browse the learning paths to see this term in context — every lesson is free to read.