Glossary · Multimodal systems
Vision Transformer (ViT)
A vision architecture that represents an image as a sequence of patch embeddings with position information and processes that sequence with transformer encoder blocks.
Why it matters
It provides a sequence-model interface for visual data, but performance and compute depend on patch size, resolution, pretraining, and inductive biases.
In practice
Keep patching and normalization consistent with training, account for position-embedding behavior at new resolutions, and compare against a suitable visual baseline on the target dataset.
Common confusion
ViT is an architecture family, not every transformer that accepts images, and its patches are not inherently semantic objects.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.