Glossary · Infrastructure & serving

Decode Phase

The iterative stage of autoregressive inference that generates new tokens one step at a time after the input prefix has been processed.

Why it matters

Decode work has different compute, memory, and scheduling behavior from prefill, so one aggregate latency number can hide the actual serving bottleneck.

In practice

Measure inter-token latency and output throughput separately, account for KV-cache occupancy, and test mixed workloads where active decodes share capacity with new prefills.

Common confusion

Decode phase is not the decoder component of an encoder-decoder model. It names the runtime generation stage.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.