Glossary · Models & inference
Self-Attention
Attention in which queries, keys, and values are derived from the same sequence representation. Scaled similarity scores are normalized and used to combine values, subject to causal, padding, local, or other masks.
Why it matters
It builds context-sensitive token representations, but the permitted attention pattern depends on the architecture.
Common confusion
Not every token can always attend to every other token. Causal and sparse models intentionally restrict connections.
Related terms
Browse the learning paths to see this term in context — every lesson is free to read.