Glossary · Multimodal systems

Multimodal Model

A model that learns from, relates, or generates more than one modality through representation, alignment, fusion, translation, or coordinated prediction.

Why it matters

Multimodal capability depends on how modalities interact, not simply on accepting several input types, and failures can occur at each representation boundary.

In practice

Document supported input and output combinations, evaluate each modality alone and together, test missing or conflicting inputs, and track preprocessing versions with the model.

Common confusion

A pipeline with separate image and text models is multimodal at the system level, but it is not necessarily one jointly trained multimodal model.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.