Glossary · Models & inference
Speculative Decoding
An inference method in which a cheaper draft process proposes several tokens and the target model scores those draft positions in parallel. In exact sampling variants, an acceptance and correction rule preserves the target model's output distribution.
Why it matters
It can reduce serial target-model decoding work when drafts are accepted, without requiring a change to the target model's trained weights.
In practice
Measure acceptance rate and end-to-end latency on real prompts, include draft-model overhead, and verify that the implementation preserves the intended decoding distribution.
Common confusion
Speculative decoding is not ordinary model routing or unverified autocomplete. Exact variants preserve the target distribution through acceptance and correction, while approximate variants may trade that guarantee for speed.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.