Glossary · Multimodal systems
Cross-Attention
Attention in which the query representation comes from one sequence or representation while keys and values come from another.
Why it matters
It gives one stream a learnable way to retrieve information from another, such as language tokens attending to visual features.
In practice
State which stream supplies queries, keys, and values, apply masks for missing or invalid positions, and inspect whether the model still performs when one modality is ablated.
Common confusion
Cross-attention is not intrinsically multimodal. It can connect two text sequences or other representations; self-attention instead derives queries, keys, and values from the same sequence representation.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.