Glossary · Models & inference

KV Cache

Stored key and value tensors from earlier positions in autoregressive generation. Reusing them avoids recomputing attention projections for the unchanged prefix at every decoding step.

Why it matters

It reduces repeated computation but consumes memory that grows with sequence length, layers, batch, and model configuration.

Common confusion

A KV cache is runtime attention state for a sequence. Prefix caching reuses eligible KV state across requests, while prompt caching is a broader provider or application reuse contract.

Related terms

Browse the learning paths to see this term in context — every lesson is free to read.