Glossary · Multimodal systems
Visual Grounding
Connecting a language expression to spatial evidence in an image or video, such as a region, object, mask, or tracked entity.
Why it matters
A fluent visual answer can be unsupported, while grounding makes the claimed referent inspectable and enables region-level evaluation.
In practice
Require a box, mask, or temporal segment with the answer, test ambiguous and absent referents, and score localization separately from language correctness.
Common confusion
Visual grounding identifies where the referenced evidence is. General image captioning can describe a scene without localizing each claim.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.