The New Engineer
PathsLessonsDashboard
Sign inStart free

Phase 06

Speech and Audio

The other half of human communication. Hear, understand, speak.

Lessons (17)

  1. 01Audio Fundamentals — Waveforms, Sampling, Fourier Transform
  2. 02Spectrograms, Mel Scale & Audio Features
  3. 03Audio Classification — From k-NN on MFCCs to AST and BEATs
  4. 04Speech Recognition (ASR) — CTC, RNN-T, Attention
  5. 05Whisper — Architecture & Fine-Tuning
  6. 06Speaker Recognition & Verification
  7. 07Text-to-Speech (TTS) — From Tacotron to F5 and Kokoro
  8. 08Voice Cloning & Voice Conversion
  9. 09Music Generation — MusicGen, Stable Audio, Suno, and the Licensing Earthquake
  10. 10Audio-Language Models — Qwen2.5-Omni, Audio Flamingo, GPT-4o Audio
  11. 11Real-Time Audio Processing
  12. 12Build a Voice Assistant Pipeline — The Phase 6 Capstone
  13. 13Neural Audio Codecs — EnCodec, SNAC, Mimi, DAC and the Semantic-Acoustic Split
  14. 14Voice Activity Detection & Turn-Taking — Silero, Cobra, and the Flush Trick
  15. 15Streaming Speech-to-Speech — Moshi, Hibiki, and Full-Duplex Dialogue
  16. 16Voice Anti-Spoofing & Audio Watermarking — ASVspoof 5, AudioSeal, WaveVerify
  17. 17Audio Evaluation — WER, MOS, UTMOS, MMAU, FAD, and the Open Leaderboards