Glossary · Multimodal systems

Vision Transformer (ViT)

A vision architecture that represents an image as a sequence of patch embeddings with position information and processes that sequence with transformer encoder blocks.

Why it matters

It provides a sequence-model interface for visual data, but performance and compute depend on patch size, resolution, pretraining, and inductive biases.

In practice

Keep patching and normalization consistent with training, account for position-embedding behavior at new resolutions, and compare against a suitable visual baseline on the target dataset.

Common confusion

ViT is an architecture family, not every transformer that accepts images, and its patches are not inherently semantic objects.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.