Phase 04
Computer Vision
From pixels to understanding across image, video, and 3D.
Lessons (28)
- 01Image Fundamentals — Pixels, Channels, Color Spaces
- 02Convolutions from Scratch
- 03CNNs — LeNet to ResNet
- 04Image Classification
- 05Transfer Learning & Fine-Tuning
- 06Object Detection — YOLO from Scratch
- 07Semantic Segmentation — U-Net
- 08Instance Segmentation — Mask R-CNN
- 09Image Generation — GANs
- 10Image Generation — Diffusion Models
- 11Stable Diffusion — Architecture & Fine-Tuning
- 12Video Understanding — Temporal Modeling
- 133D Vision — Point Clouds & NeRFs
- 14Vision Transformers (ViT)
- 15Real-Time Vision — Edge Deployment
- 16Build a Complete Vision Pipeline — Capstone
- 17Self-Supervised Vision — SimCLR, DINO, MAE
- 18Open-Vocabulary Vision — CLIP
- 19OCR & Document Understanding
- 20Image Retrieval & Metric Learning
- 21Keypoint Detection & Pose Estimation
- 223D Gaussian Splatting from Scratch
- 23Diffusion Transformers & Rectified Flow
- 24SAM 3 & Open-Vocabulary Segmentation
- 25Vision-Language Models — The ViT-MLP-LLM Pattern
- 26Monocular Depth & Geometry Estimation
- 27Multi-Object Tracking & Video Memory
- 28World Models & Video Diffusion