Glossary · Multimodal systems

Vision-Language Model (VLM)

A model that learns relationships between, or jointly processes, visual and language representations for tasks such as retrieval, description, question answering, or grounded generation.

Why it matters

VLM performance depends on the visual encoder, language component, connection mechanism, training data, and resolution policy rather than one generic capability label.

In practice

Evaluate text-only and vision-only controls, vary image resolution and layout, require evidence localization where possible, and report failures by visual skill and language.

Common confusion

Accepting an image does not prove the model uses it correctly, and a VLM is not necessarily able to generate images.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.