Glossary · Multimodal systems

Patch Embedding

A learned projection that converts an image patch into a fixed-width vector used as one element of a transformer input sequence.

Why it matters

It creates the interface between a spatial image grid and a sequence model, with patch size controlling token count and retained local detail.

In practice

Record patch and image dimensions, handle padding or resizing explicitly, add position information, and measure how resolution changes affect both accuracy and token cost.

Common confusion

A patch embedding is the vector representation of a patch, not a semantic object detector or a guarantee that patch boundaries match visual entities.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.