Glossary · Multimodal systems
Vision-Language Model (VLM)
A model that learns relationships between, or jointly processes, visual and language representations for tasks such as retrieval, description, question answering, or grounded generation.
Why it matters
VLM performance depends on the visual encoder, language component, connection mechanism, training data, and resolution policy rather than one generic capability label.
In practice
Evaluate text-only and vision-only controls, vary image resolution and layout, require evidence localization where possible, and report failures by visual skill and language.
Common confusion
Accepting an image does not prove the model uses it correctly, and a VLM is not necessarily able to generate images.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.