Glossary · Models & inference

Transformer

A neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization. Encoder, decoder, and encoder-decoder variants use different masks and information flows.

Why it matters

Training can process many sequence positions in parallel, while autoregressive generation still produces outputs step by step.

Common confusion

Self-attention does not imply unrestricted all-to-all attention in every transformer.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.