Glossary · Multimodal systems

Cross-Attention

Attention in which the query representation comes from one sequence or representation while keys and values come from another.

Why it matters

It gives one stream a learnable way to retrieve information from another, such as language tokens attending to visual features.

In practice

State which stream supplies queries, keys, and values, apply masks for missing or invalid positions, and inspect whether the model still performs when one modality is ablated.

Common confusion

Cross-attention is not intrinsically multimodal. It can connect two text sequences or other representations; self-attention instead derives queries, keys, and values from the same sequence representation.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.