Phase 07
Transformers Deep Dive
The architecture that changed everything. Understand every layer.
Lessons (16)
- 01Why Transformers — The Problems with RNNs
- 02Self-Attention from Scratch
- 03Multi-Head Attention
- 04Positional Encoding — Sinusoidal, RoPE, ALiBi
- 05The Full Transformer — Encoder + Decoder
- 06BERT — Masked Language Modeling
- 07GPT — Causal Language Modeling
- 08T5, BART — Encoder-Decoder Models
- 09Vision Transformers (ViT)
- 10Audio Transformers — Whisper Architecture
- 11Mixture of Experts (MoE)
- 12KV Cache, Flash Attention & Inference Optimization
- 13Scaling Laws
- 14Build a Transformer from Scratch — The Capstone
- 15Attention Variants — Sliding Window, Sparse, Differential
- 16Speculative Decoding — Draft, Verify, Repeat