Glossary · Models & inference
KV Cache
Stored key and value tensors from earlier positions in autoregressive generation. Reusing them avoids recomputing attention projections for the unchanged prefix at every decoding step.
Why it matters
It reduces repeated computation but consumes memory that grows with sequence length, layers, batch, and model configuration.
Common confusion
A KV cache is runtime attention state for a sequence. Prefix caching reuses eligible KV state across requests, while prompt caching is a broader provider or application reuse contract.
Related terms
Browse the learning paths to see this term in context — every lesson is free to read.