The New Engineer
PathsLessonsDashboard
Sign inStart free

Phase 07

Transformers Deep Dive

The architecture that changed everything. Understand every layer.

Lessons (16)

  1. 01Why Transformers — The Problems with RNNs
  2. 02Self-Attention from Scratch
  3. 03Multi-Head Attention
  4. 04Positional Encoding — Sinusoidal, RoPE, ALiBi
  5. 05The Full Transformer — Encoder + Decoder
  6. 06BERT — Masked Language Modeling
  7. 07GPT — Causal Language Modeling
  8. 08T5, BART — Encoder-Decoder Models
  9. 09Vision Transformers (ViT)
  10. 10Audio Transformers — Whisper Architecture
  11. 11Mixture of Experts (MoE)
  12. 12KV Cache, Flash Attention & Inference Optimization
  13. 13Scaling Laws
  14. 14Build a Transformer from Scratch — The Capstone
  15. 15Attention Variants — Sliding Window, Sparse, Differential
  16. 16Speculative Decoding — Draft, Verify, Repeat