Phase 17

Infrastructure and Production

Ship AI to the real world. Scale, monitor, optimize.

Lessons (28)

  1. 01Managed LLM Platforms — Bedrock, Vertex AI, Azure OpenAI
  2. 02Inference Platform Economics — Fireworks, Together, Baseten, Modal, Replicate, Anyscale
  3. 03GPU Autoscaling on Kubernetes — Karpenter, KAI Scheduler, Gang Scheduling
  4. 04Serving Engine Internals — PagedAttention, Continuous Batching, Chunked Prefill
  5. 05EAGLE-3 Speculative Decoding in Production
  6. 06Prefix-Cache Serving — RadixAttention and KV Reuse
  7. 07Hardware-Specialized Inference Compilation — FP8 and NVFP4 on Blackwell
  8. 08Inference Metrics — TTFT, TPOT, ITL, Goodput, P99
  9. 09Production Quantization — AWQ, GPTQ, GGUF K-quants, FP8, MXFP4/NVFP4
  10. 10Cold Start Mitigation for Serverless LLMs
  11. 11Multi-Region LLM Serving and KV Cache Locality
  12. 12Edge Inference — Apple Neural Engine, Qualcomm Hexagon, WebGPU/WebLLM, Jetson
  13. 13LLM Observability Stack Selection
  14. 14Prompt Caching and Semantic Caching Economics
  15. 15Batch APIs — the 50% Discount as Industry Standard
  16. 16Model Routing as a Cost-Reduction Primitive
  17. 17Disaggregated Prefill/Decode — NVIDIA Dynamo and llm-d
  18. 18Production Serving Stack — KV Offloading and Cache-Aware Routing
  19. 19AI Gateways — LiteLLM, Portkey, Kong AI Gateway, Bifrost
  20. 20Shadow Traffic, Canary Rollout, and Progressive Deployment for LLMs
  21. 21A/B Testing LLM Features — GrowthBook, Statsig, and the Vibes Problem
  22. 22Load Testing LLM APIs — Why k6 and Locust Lie
  23. 23SRE for AI — Multi-Agent Incident Response, Runbooks, Predictive Detection
  24. 24Chaos Engineering for LLM Production
  25. 25Security — Secrets, API Key Rotation, Audit Logs, Guardrails
  26. 26Compliance — SOC 2, HIPAA, GDPR, PCI-DSS, EU AI Act, ISO 42001
  27. 27FinOps for LLMs — Unit Economics and Multi-Tenant Attribution
  28. 28Self-Hosted Serving Selection — Matching Engine to Hardware and Scale

Learning paths covering this phase