Phase 14

Agent Engineering

The core of modern AI engineering. Build agents from first principles.

Lessons (54)

  1. 01The Agent Loop: Observe, Think, Act
  2. 02ReWOO and Plan-and-Execute: Decoupled Planning
  3. 03Reflexion: Verbal Reinforcement Learning
  4. 04Tree of Thoughts and LATS: Deliberate Search
  5. 05Self-Refine and CRITIC: Iterative Output Improvement
  6. 06Tool Use and Function Calling
  7. 07Agent Memory — Virtual Context and Memory Paging
  8. 08Memory Blocks and Sleep-Time Compute
  9. 09Hybrid Memory: Vector + Graph + KV
  10. 10Skill Libraries and Lifelong Learning (Voyager)
  11. 11Planning with HTN and Evolutionary Search
  12. 12Anthropic's Workflow Patterns: Simple Over Complex
  13. 13Stateful Graph Orchestration — Durable Execution and Checkpoints
  14. 14The Actor Model for Agents — Async Messages and Typed Runtimes
  15. 15Role-Based Agent Teams — Roles, Tasks, Processes
  16. 16OpenAI Agents SDK: Handoffs, Guardrails, Tracing
  17. 17The Harness as a Library — Subagents and Session Store
  18. 18Production Agent Runtimes — Fast Instantiation and Typed Workflows
  19. 19Benchmarks: SWE-bench, GAIA, AgentBench
  20. 20Benchmarks: WebArena and OSWorld
  21. 21Computer Use: Claude, OpenAI CUA, Gemini
  22. 22Voice Agents: Pipecat and LiveKit
  23. 23OpenTelemetry GenAI Semantic Conventions
  24. 24Agent Observability: Langfuse, Phoenix, Opik
  25. 25Multi-Agent Debate and Collaboration
  26. 26Failure Modes: Why Agents Break
  27. 27Prompt Injection and the PVE Defense
  28. 28Orchestration Patterns: Supervisor, Swarm, Hierarchical
  29. 29Production Runtimes: Queue, Event, Cron
  30. 30Eval-Driven Agent Development
  31. 31Agent Workbench Engineering: Why Capable Models Still Fail
  32. 32The Minimal Agent Workbench
  33. 33Agent Instructions as Executable Constraints
  34. 34Repo Memory and Durable State
  35. 35Initialization Scripts for Agents
  36. 36Scope Contracts and Task Boundaries
  37. 37Runtime Feedback Loops
  38. 38Verification Gates
  39. 39Reviewer Agent: Separate Builder from Marker
  40. 40Multi-Session Handoff
  41. 41The Workbench on a Real Repo
  42. 42Capstone: Ship a Reusable Agent Workbench Pack
  43. 43Frame the Task Before the Agent Writes Code
  44. 44Build an Evidence-Backed Execution Plan
  45. 45Delegate Agent Work with Isolation and Merge Contracts
  46. 46Turn Every Agent Correction into a System Improvement
  47. 47Define the Outcome Before You Choose the Output
  48. 48Discover the Workflow People Actually Perform
  49. 49Map Assumptions and Resolve the Riskiest One First
  50. 50Choose the Smallest Slice That Can Change the Decision
  51. 51Write Specifications That Preserve Judgment
  52. 52Design Success Metrics Before the Result Exists
  53. 53Choose Prototype, Pilot, or Production Deliberately
  54. 54Build a Feedback Ratchet with Ownership and Retirement

Learning paths covering this phase