The New Engineer
PathsLessonsDashboard
Sign inStart free

Phase 18

Ethics Safety Alignment

Build AI that helps humanity. Not optional.

Lessons (30)

  1. 01Instruction-Following as Alignment Signal
  2. 02Reward Hacking and Goodhart's Law
  3. 03The Direct Preference Optimization Family
  4. 04Sycophancy as RLHF Amplification
  5. 05Constitutional AI and RLAIF
  6. 06Mesa-Optimization and Deceptive Alignment
  7. 07Sleeper Agents — Persistent Deception
  8. 08In-Context Scheming in Frontier Models
  9. 09Alignment Faking
  10. 10AI Control — Safety Despite Subversion
  11. 11Scalable Oversight and Weak-to-Strong Generalization
  12. 12Red-Teaming: PAIR and Automated Attacks
  13. 13Many-Shot Jailbreaking
  14. 14ASCII Art and Visual Jailbreaks
  15. 15Indirect Prompt Injection — Production Attack Surface
  16. 16Red-Team Tooling — Garak, Llama Guard, PyRIT
  17. 17WMDP and Dual-Use Capability Evaluation
  18. 18Frontier Safety Frameworks — RSP, PF, FSF
  19. 19Anthropic's Model Welfare Program
  20. 20Bias and Representational Harm in LLMs
  21. 21Fairness Criteria — Group, Individual, Counterfactual
  22. 22Differential Privacy for LLMs
  23. 23Watermarking — SynthID, Stable Signature, C2PA
  24. 24Regulatory Frameworks — EU, US, UK, Korea
  25. 25EchoLeak and the Emergence of CVEs for AI
  26. 26Model, System, and Dataset Cards
  27. 27Data Provenance and Training-Data Governance
  28. 28Alignment Research Ecosystem — MATS, Redwood, Apollo, METR
  29. 29Moderation Systems — OpenAI, Perspective, Llama Guard
  30. 30Dual-Use Risk — Cyber, Bio, Chem, Nuclear Uplift