Glossary · Evaluation & safety

Pass@k

Across a task set, the fraction of tasks for which at least one of k sampled candidates passes a defined correctness test.

Why it matters

It measures the value of sampling several attempts for tasks such as code generation where an automatic verifier can check each candidate.

In practice

Generate candidates independently under a fixed configuration, run the same isolated tests on each, and report k with the sampling and estimator details.

Common confusion

Pass@k is not single-attempt accuracy, and a higher score can reflect a larger attempt budget rather than a better first answer.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.