The New Engineer
PathsLessonsDashboard
Sign inStart free

Phase 09

Reinforcement Learning

Agents that learn by doing. The foundation of RLHF.

Lessons (12)

  1. 01MDPs, States, Actions & Rewards
  2. 02Dynamic Programming — Policy Iteration & Value Iteration
  3. 03Monte Carlo Methods — Learning from Complete Episodes
  4. 04Temporal Difference — Q-Learning & SARSA
  5. 05Deep Q-Networks (DQN)
  6. 06Policy Gradient — REINFORCE from Scratch
  7. 07Actor-Critic — A2C and A3C
  8. 08Proximal Policy Optimization (PPO)
  9. 09Reward Modeling & RLHF
  10. 10Multi-Agent RL
  11. 11Sim-to-Real Transfer
  12. 12RL for Games — AlphaZero, MuZero, and the LLM-Reasoning Era