Glossary · Infrastructure & serving
Decode Phase
The iterative stage of autoregressive inference that generates new tokens one step at a time after the input prefix has been processed.
Why it matters
Decode work has different compute, memory, and scheduling behavior from prefill, so one aggregate latency number can hide the actual serving bottleneck.
In practice
Measure inter-token latency and output throughput separately, account for KV-cache occupancy, and test mixed workloads where active decodes share capacity with new prefills.
Common confusion
Decode phase is not the decoder component of an encoder-decoder model. It names the runtime generation stage.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.