Text-to-speech models
A text-to-speech model synthesizes speech directly from text.
Text-to-speech synthesizes speech from text. Tacotron 2 first predicts mel spectrograms and then uses a modified WaveNet vocoder to synthesize waveforms from them.
Tacotron 2 maps character embeddings to mel-scale spectrograms. A modified WaveNet vocoder synthesizes time-domain waveforms from those spectrograms.
Tacotron synthesized speech directly from characters. FastSpeech generates mel spectrograms in parallel. VITS supports single-stage training and parallel sampling. Its paper notes that one text input can be spoken with different pitches and rhythms.
Tacotron can synthesize speech directly from characters.
Follow one sentence through a Tacotron 2 style system.
- 1 · encodeConvert characters into embeddings.
- 2 · predictMap the character embeddings to a mel-scale spectrogram.
- 3 · synthesizeCondition a vocoder on that spectrogram.
- 4 · playReturn the time-domain waveform as speech audio.
VITS supports single-stage training and parallel sampling.
| Who | What they ask | What it works with |
|---|---|---|
| Accessibility team | “How should this passage sound when read aloud?” | Text and synthesized audio |
| Speech engineer | “Where are skipped or repeated words introduced?” | Acoustic-model outputs and durations |
| Runtime team | “Which stage limits synthesis speed?” | Spectrogram generator and vocoder timings |
- An end-to-end system can synthesize speech directly from characters.
- A mel spectrogram can act as the intermediate representation before waveform synthesis.
- FastSpeech generates mel spectrograms in parallel.
- VITS supports single-stage training and parallel sampling.
- A typical text-to-speech pipeline has multiple stages.
- One text input can be spoken in multiple ways with different pitches and rhythms.
- An autoregressive WaveNet predicts each audio sample from previous samples.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperTacotron: Towards End-to-End Speech Synthesis, Wang and colleagues · read 28 Sept 2026
- paperNatural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions, Shen and colleagues · read 28 Sept 2026
- paperFastSpeech: Fast, Robust and Controllable Text to Speech, Ren and colleagues · read 28 Sept 2026
- paperConditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech, Kim, Kong and Son · read 28 Sept 2026
- paperWaveNet: A Generative Model for Raw Audio, van den Oord and colleagues · read 28 Sept 2026