Phase 14
Agent Engineering
The core of modern AI engineering. Build agents from first principles.
Lessons (54)
- 01The Agent Loop: Observe, Think, Act
- 02ReWOO and Plan-and-Execute: Decoupled Planning
- 03Reflexion: Verbal Reinforcement Learning
- 04Tree of Thoughts and LATS: Deliberate Search
- 05Self-Refine and CRITIC: Iterative Output Improvement
- 06Tool Use and Function Calling
- 07Agent Memory — Virtual Context and Memory Paging
- 08Memory Blocks and Sleep-Time Compute
- 09Hybrid Memory: Vector + Graph + KV
- 10Skill Libraries and Lifelong Learning (Voyager)
- 11Planning with HTN and Evolutionary Search
- 12Anthropic's Workflow Patterns: Simple Over Complex
- 13Stateful Graph Orchestration — Durable Execution and Checkpoints
- 14The Actor Model for Agents — Async Messages and Typed Runtimes
- 15Role-Based Agent Teams — Roles, Tasks, Processes
- 16OpenAI Agents SDK: Handoffs, Guardrails, Tracing
- 17The Harness as a Library — Subagents and Session Store
- 18Production Agent Runtimes — Fast Instantiation and Typed Workflows
- 19Benchmarks: SWE-bench, GAIA, AgentBench
- 20Benchmarks: WebArena and OSWorld
- 21Computer Use: Claude, OpenAI CUA, Gemini
- 22Voice Agents: Pipecat and LiveKit
- 23OpenTelemetry GenAI Semantic Conventions
- 24Agent Observability: Langfuse, Phoenix, Opik
- 25Multi-Agent Debate and Collaboration
- 26Failure Modes: Why Agents Break
- 27Prompt Injection and the PVE Defense
- 28Orchestration Patterns: Supervisor, Swarm, Hierarchical
- 29Production Runtimes: Queue, Event, Cron
- 30Eval-Driven Agent Development
- 31Agent Workbench Engineering: Why Capable Models Still Fail
- 32The Minimal Agent Workbench
- 33Agent Instructions as Executable Constraints
- 34Repo Memory and Durable State
- 35Initialization Scripts for Agents
- 36Scope Contracts and Task Boundaries
- 37Runtime Feedback Loops
- 38Verification Gates
- 39Reviewer Agent: Separate Builder from Marker
- 40Multi-Session Handoff
- 41The Workbench on a Real Repo
- 42Capstone: Ship a Reusable Agent Workbench Pack
- 43Frame the Task Before the Agent Writes Code
- 44Build an Evidence-Backed Execution Plan
- 45Delegate Agent Work with Isolation and Merge Contracts
- 46Turn Every Agent Correction into a System Improvement
- 47Define the Outcome Before You Choose the Output
- 48Discover the Workflow People Actually Perform
- 49Map Assumptions and Resolve the Riskiest One First
- 50Choose the Smallest Slice That Can Change the Decision
- 51Write Specifications That Preserve Judgment
- 52Design Success Metrics Before the Result Exists
- 53Choose Prototype, Pilot, or Production Deliberately
- 54Build a Feedback Ratchet with Ownership and Retirement