The New Engineer
PathsLessonsDashboard
Sign inStart free

Phase 12

Multimodal AI

Models that see, hear, read, and reason across modalities.

Lessons (25)

  1. 01Vision Transformers and the Patch-Token Primitive
  2. 02CLIP and Contrastive Vision-Language Pretraining
  3. 03From CLIP to BLIP-2 — Q-Former as Modality Bridge
  4. 04Flamingo and Gated Cross-Attention for Few-Shot VLMs
  5. 05LLaVA and Visual Instruction Tuning
  6. 06Any-Resolution Vision: Patch-n'-Pack and NaFlex
  7. 07Open-Weight VLM Recipes: What Actually Matters
  8. 08LLaVA-OneVision: Single-Image, Multi-Image, Video in One Model
  9. 09Qwen-VL Family and Dynamic-FPS Video
  10. 10InternVL3: Native Multimodal Pretraining
  11. 11Chameleon and Early-Fusion Token-Only Multimodal Models
  12. 12Emu3: Next-Token Prediction for Image and Video Generation
  13. 13Transfusion: Autoregressive Text + Diffusion Image in One Transformer
  14. 14Show-o and Discrete-Diffusion Unified Models
  15. 15Janus-Pro: Decoupled Encoders for Unified Multimodal Models
  16. 16MIO and Any-to-Any Streaming Multimodal Models
  17. 17Video-Language Models: Temporal Tokens and Grounding
  18. 18Long-Video Understanding at Million-Token Context
  19. 19Audio-Language Models: the Whisper to Audio Flamingo 3 Arc
  20. 20Omni Models: Qwen2.5-Omni and the Thinker-Talker Split
  21. 21Embodied VLAs: RT-2, OpenVLA, π0, GR00T
  22. 22Document and Diagram Understanding
  23. 23ColPali and Vision-Native Document RAG
  24. 24Multimodal RAG and Cross-Modal Retrieval
  25. 25Multimodal Agents and Computer-Use (Capstone)