Glossary · Multimodal systems

Visual Grounding

Connecting a language expression to spatial evidence in an image or video, such as a region, object, mask, or tracked entity.

Why it matters

A fluent visual answer can be unsupported, while grounding makes the claimed referent inspectable and enables region-level evaluation.

In practice

Require a box, mask, or temporal segment with the answer, test ambiguous and absent referents, and score localization separately from language correctness.

Common confusion

Visual grounding identifies where the referenced evidence is. General image captioning can describe a scene without localizing each claim.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.