Glossary · Infrastructure & serving
Paged KV Cache
A KV-cache memory manager that stores attention state in fixed-size blocks and maps logical sequence positions to physical blocks instead of requiring one contiguous allocation per sequence.
Why it matters
Variable sequence lengths create fragmentation and unpredictable growth, so block-based allocation can improve usable memory and enable flexible sharing.
In practice
Select block size from workload measurements, track allocation and eviction, isolate state between requests, and test cancellation and prefix sharing under memory pressure.
Common confusion
Paged KV cache manages runtime attention-state memory. It does not move model parameters to disk or extend the model's trained context limit.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.