A decoding method that samples from the smallest set of next-token candidates whose cumulative probability reaches a chosen threshold.
Why it matters
The candidate-set size adapts to the distribution, retaining more options when uncertainty is broad and fewer when probability is concentrated.
In practice
Evaluate the threshold with temperature and stop settings held constant, and record the complete decoding configuration with every result.
Common confusion
Top-p is a probability-mass threshold, while top-k always keeps a fixed maximum number of candidates.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.