Glossary · Infrastructure & serving

Prefill

Also known as: Prefill Phase

The initial inference stage that processes all supplied input tokens to produce their representations and the attention state required for subsequent autoregressive generation.

Why it matters

Prompt shape, queueing, and cache reuse affect prefill cost, and prefill competes differently for compute than decode, so it strongly influences startup latency and serving schedules.

In practice

Record prompt tokens and prefill latency, separate queue time from execution time, compare cached and uncached prefixes, and test long prompts beside active decode traffic.

Common confusion

Prefill is the runtime prompt-processing stage, not the first generated token itself. The first token appears only after prefill and any queueing complete.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.