Glossary · Multimodal systems
Patch Embedding
A learned projection that converts an image patch into a fixed-width vector used as one element of a transformer input sequence.
Why it matters
It creates the interface between a spatial image grid and a sequence model, with patch size controlling token count and retained local detail.
In practice
Record patch and image dimensions, handle padding or resizing explicitly, add position information, and measure how resolution changes affect both accuracy and token cost.
Common confusion
A patch embedding is the vector representation of a patch, not a semantic object detector or a guarantee that patch boundaries match visual entities.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.