Glossary · Infrastructure & serving
FlashAttention
An exact attention algorithm that tiles the computation to reduce transfers between accelerator memory levels while avoiding materialization of the full attention matrix in high-bandwidth memory.
Why it matters
Attention can be limited by memory movement rather than arithmetic, especially for long sequences, so an IO-aware kernel can improve usable speed and memory efficiency.
In practice
Use a kernel supported by the model's shapes, masks, dtype, and hardware, verify numerical tolerance, and benchmark end-to-end latency rather than quoting a paper result as a fixed multiplier.
Common confusion
FlashAttention changes how attention is computed, not the mathematical attention result it targets. It is separate from KV caching and quantization.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.