The initial inference stage that processes all supplied input tokens to produce their representations and the attention state required for subsequent autoregressive generation.
Why it matters
Prompt shape, queueing, and cache reuse affect prefill cost, and prefill competes differently for compute than decode, so it strongly influences startup latency and serving schedules.
In practice
Record prompt tokens and prefill latency, separate queue time from execution time, compare cached and uncached prefixes, and test long prompts beside active decode traffic.
Common confusion
Prefill is the runtime prompt-processing stage, not the first generated token itself. The first token appears only after prefill and any queueing complete.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.