Glossary · Multimodal systems

Image Token

A model-specific visual unit represented as a vector or discrete code, commonly derived from an image patch, region, or learned visual-codebook entry.

Why it matters

Turning visual input into a sequence lets transformer-style components process images together with text or other tokenized modalities.

In practice

Document whether tokens are continuous patches or discrete codes, preserve spatial position, test resolution and aspect-ratio changes, and count visual tokens in the model's input budget.

Common confusion

An image token is not necessarily one pixel, one object, or one fixed physical area. Its scope follows the visual encoder or tokenizer.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.