Glossary · Models & inference
Transformer
A neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization. Encoder, decoder, and encoder-decoder variants use different masks and information flows.
Why it matters
Training can process many sequence positions in parallel, while autoregressive generation still produces outputs step by step.
Common confusion
Self-attention does not imply unrestricted all-to-all attention in every transformer.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.