All paths

AI Evaluation and Reliability

Career route

Measure model and agent behavior, expose failure modes, instrument the runtime, and build release and incident controls around evidence.

~13h12 lessons

1. Common Core

5 lessons

Define behavior, evaluation layers, failure modes, and trace semantics.

You'll build: A layered evaluation plan with a failure taxonomy, trace schema, and regression thresholds.

  1. Model Evaluation~2h
  2. Evaluation & Testing LLM Applications~1h
  3. Eval-Driven Agent Development~1h
  4. Failure Modes: Why Agents Break~1h
  5. OpenTelemetry GenAI Semantic Conventions~1h

2. Role Practice

4 lessons

Measure inference health, production behavior, experiments, and capacity limits.

You'll build: An observability and load report that connects service metrics with behavior and user outcomes.

  1. Inference Metrics — TTFT, TPOT, ITL, Goodput, P99~1h
  2. LLM Observability Stack Selection~1h
  3. A/B Testing LLM Features — GrowthBook, Statsig, and the Vibes Problem~1h
  4. Load Testing LLM APIs — Why k6 and Locust Lie~1h

3. Proof Project

2 lessons

Prove release and recovery behavior under controlled failure.

You'll build: A regression and reliability gate with canary thresholds, rollback evidence, and a chaos result.

  1. Shadow Traffic, Canary Rollout, and Progressive Deployment for LLMs~1h
  2. Chaos Engineering for LLM Production~1h

4. Interview and Readiness Evidence

1 lessons

Show how objectives, incidents, and corrective actions form an operating system for reliability.

You'll build: A reliability case study with SLOs, failure evidence, response decisions, and verified corrective actions.

  1. SRE for AI — Multi-Agent Incident Response, Runbooks, Predictive Detection~1h