Phase 09
Reinforcement Learning
Agents that learn by doing. The foundation of RLHF.
Lessons (12)
- 01MDPs, States, Actions & Rewards
- 02Dynamic Programming — Policy Iteration & Value Iteration
- 03Monte Carlo Methods — Learning from Complete Episodes
- 04Temporal Difference — Q-Learning & SARSA
- 05Deep Q-Networks (DQN)
- 06Policy Gradient — REINFORCE from Scratch
- 07Actor-Critic — A2C and A3C
- 08Proximal Policy Optimization (PPO)
- 09Reward Modeling & RLHF
- 10Multi-Agent RL
- 11Sim-to-Real Transfer
- 12RL for Games — AlphaZero, MuZero, and the LLM-Reasoning Era