Speech Synthesis and Recognition (TTS & ASR)
Speech modeling runs in two mirrored directions: TTS expands text down to a waveform, ASR compresses a waveform up to words. This topic traces both from hand-glued HMM modules to end-to-end neural nets - acoustic models, vocoders, CTC/Transducer/seq2seq heads, self-supervised encoders, and RVQ codec tokens - and shows how 2024-2026 native-audio LLMs fold the two pipelines into one.
01.The Problem: Speech Is Numbers Too Big to Handle Directly
Open a recording of someone saying "hello world" and look at what a computer actually sees.
A phone samples sound 16,000 times per second. Two seconds of speech = 32,000 floating-point numbers. Nobody trains a network to spit out 32,000 precise values one by one.
Both speech tasks are really about squashing that river of numbers into symbols and back:
- TTS (text-to-speech, synthesis): text → a waveform. Expanding 11 letters into 32,000 numbers that sound human.
- ASR (automatic speech recognition, transcription): a waveform → text. Compressing 32,000 numbers into 11 letters, correctly, at speed.
They are the same sequence-to-sequence problem pointed in opposite directions.
Carry one analogy: two embassies, Text-land and Sound-land, exchanging mail across a border.
- The outbound ambassador (TTS) turns a letter into spoken sound.
- The inbound ambassador (ASR) turns heard sound into a letter.
- Historically each embassy had its own bizarre internal bureaucracy of hand-designed departments. The last decade of speech AI is the story of replacing those departments with neural nets — and the current chapter is the two embassies merging into one bilingual brain.
Two scoreboards matter throughout. ASR is judged by WER (Word Error Rate) — align the transcript against the truth and count substitutions + deletions + insertions over the reference length (100 correct words with 7 wrong → 7% WER) — plus latency for live use. TTS is judged by MOS (Mean Opinion Score): humans rate naturalness/intelligibility 1-5, averaged — and by speaker similarity when cloning a specific voice.
Two Mirrored Pipelines
Two Mirrored Pipelines
TTS expands text down to a waveform through an acoustic model plus a vocoder; ASR compresses audio up to text through an encoder plus a CTC/RNN-T/seq2seq decoder. Modern neural-codec models blur the line by treating audio as tokens a language model can both read and write.
Unlock Topic #193: Speech Synthesis and Recognition (TTS & ASR)
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?