Glossary · Infrastructure & serving

FlashAttention

An exact attention algorithm that tiles the computation to reduce transfers between accelerator memory levels while avoiding materialization of the full attention matrix in high-bandwidth memory.

Why it matters

Attention can be limited by memory movement rather than arithmetic, especially for long sequences, so an IO-aware kernel can improve usable speed and memory efficiency.

In practice

Use a kernel supported by the model's shapes, masks, dtype, and hardware, verify numerical tolerance, and benchmark end-to-end latency rather than quoting a paper result as a fixed multiplier.

Common confusion

FlashAttention changes how attention is computed, not the mathematical attention result it targets. It is separate from KV caching and quantization.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.