The New Engineer
PathsLessonsDashboard
Sign inStart free

Phase 04

Computer Vision

From pixels to understanding across image, video, and 3D.

Lessons (28)

  1. 01Image Fundamentals — Pixels, Channels, Color Spaces
  2. 02Convolutions from Scratch
  3. 03CNNs — LeNet to ResNet
  4. 04Image Classification
  5. 05Transfer Learning & Fine-Tuning
  6. 06Object Detection — YOLO from Scratch
  7. 07Semantic Segmentation — U-Net
  8. 08Instance Segmentation — Mask R-CNN
  9. 09Image Generation — GANs
  10. 10Image Generation — Diffusion Models
  11. 11Stable Diffusion — Architecture & Fine-Tuning
  12. 12Video Understanding — Temporal Modeling
  13. 133D Vision — Point Clouds & NeRFs
  14. 14Vision Transformers (ViT)
  15. 15Real-Time Vision — Edge Deployment
  16. 16Build a Complete Vision Pipeline — Capstone
  17. 17Self-Supervised Vision — SimCLR, DINO, MAE
  18. 18Open-Vocabulary Vision — CLIP
  19. 19OCR & Document Understanding
  20. 20Image Retrieval & Metric Learning
  21. 21Keypoint Detection & Pose Estimation
  22. 223D Gaussian Splatting from Scratch
  23. 23Diffusion Transformers & Rectified Flow
  24. 24SAM 3 & Open-Vocabulary Segmentation
  25. 25Vision-Language Models — The ViT-MLP-LLM Pattern
  26. 26Monocular Depth & Geometry Estimation
  27. 27Multi-Object Tracking & Video Memory
  28. 28World Models & Video Diffusion