Concepts

Text-to-speech models

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

A text-to-speech model synthesizes speech directly from text.

1 · What it is

Text-to-speech synthesizes speech from text. Tacotron 2 first predicts mel spectrograms and then uses a modified WaveNet vocoder to synthesize waveforms from them.

Tacotron 2 maps character embeddings to mel-scale spectrograms. A modified WaveNet vocoder synthesizes time-domain waveforms from those spectrograms.

Tacotron synthesized speech directly from characters. FastSpeech generates mel spectrograms in parallel. VITS supports single-stage training and parallel sampling. Its paper notes that one text input can be spoken with different pitches and rhythms.

2 · Why it exists

Tacotron can synthesize speech directly from characters.

Text to acousticsTacotron 2 maps character embeddings to mel-scale spectrograms.
Acoustics to audioIts modified WaveNet vocoder synthesizes waveforms from those spectrograms.
Timing variesOne text input can be spoken with different pitches and rhythms.
3 · How it works

Follow one sentence through a Tacotron 2 style system.

Tacotron 2 predicts mel spectrograms and uses a modified WaveNet vocoder to synthesize waveforms from them.
  1. 1 · encodeConvert characters into embeddings.
  2. 2 · predictMap the character embeddings to a mel-scale spectrogram.
  3. 3 · synthesizeCondition a vocoder on that spectrogram.
  4. 4 · playReturn the time-domain waveform as speech audio.

VITS supports single-stage training and parallel sampling.

4 · Where it's used
WhoWhat they askWhat it works with
Accessibility team“How should this passage sound when read aloud?”Text and synthesized audio
Speech engineer“Where are skipped or repeated words introduced?”Acoustic-model outputs and durations
Runtime team“Which stage limits synthesis speed?”Spectrogram generator and vocoder timings
5 · What it solves, and what it doesn't
solves
  • An end-to-end system can synthesize speech directly from characters.
  • A mel spectrogram can act as the intermediate representation before waveform synthesis.
  • FastSpeech generates mel spectrograms in parallel.
  • VITS supports single-stage training and parallel sampling.
doesn't solve
  • A typical text-to-speech pipeline has multiple stages.
  • One text input can be spoken in multiple ways with different pitches and rhythms.
  • An autoregressive WaveNet predicts each audio sample from previous samples.