Phase 10
LLMs from Scratch
Build, train, and understand large language models.
Lessons (24)
- 01Tokenizers: BPE, WordPiece, SentencePiece
- 02Building a Tokenizer from Scratch
- 03Data Pipelines for Pre-Training
- 04Pre-Training a Mini GPT (124M Parameters)
- 05Scaling: Distributed Training, FSDP, DeepSpeed
- 06Instruction Tuning (SFT)
- 07RLHF: Reward Model + PPO
- 08DPO: Direct Preference Optimization
- 09Constitutional AI and Self-Improvement
- 10Evaluation: Benchmarks, Evals, LM Harness
- 11Quantization: Making Models Fit
- 12Inference Optimization
- 13Building a Complete LLM Pipeline
- 14Open Models: Architecture Walkthroughs
- 15Speculative Decoding and EAGLE-3
- 16Differential Attention (V2)
- 17Native Sparse Attention (DeepSeek NSA)
- 18Multi-Token Prediction (MTP)
- 19DualPipe Parallelism
- 20DeepSeek-V3 Architecture Walkthrough
- 21Jamba — Hybrid SSM-Transformer
- 22Async and Hogwild! Inference
- 25Speculative Decoding and EAGLE
- 34Gradient Checkpointing and Activation Recomputation