Glossary · Models & inference

Speculative Decoding

An inference method in which a cheaper draft process proposes several tokens and the target model scores those draft positions in parallel. In exact sampling variants, an acceptance and correction rule preserves the target model's output distribution.

Why it matters

It can reduce serial target-model decoding work when drafts are accepted, without requiring a change to the target model's trained weights.

In practice

Measure acceptance rate and end-to-end latency on real prompts, include draft-model overhead, and verify that the implementation preserves the intended decoding distribution.

Common confusion

Speculative decoding is not ordinary model routing or unverified autocomplete. Exact variants preserve the target distribution through acceptance and correction, while approximate variants may trade that guarantee for speed.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.