Concepts

Speech-to-text models

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

A speech-to-text model can transcribe a speech utterance into written characters.

1 · What it is

Speech-to-text turns audio into a written sequence. In sequence transduction, the alignment between discrete input and output sequences can be unknown.

Listen, Attend and Spell uses two components. A pyramidal recurrent listener accepts filter-bank spectra. An attention-based recurrent speller emits characters. The complete recognizer learns its components jointly.

Other systems use different training routes. Deep Speech used end-to-end deep learning without a phoneme dictionary. Whisper training used multilingual and multitask supervision. wav2vec 2.0 masked speech input in latent space and solved a contrastive task before transcript fine-tuning.

2 · Why it exists

Speech transduction can involve an unknown input-output alignment.

Find alignmentThe alignment between input and output sequences can be unknown.
Handle acousticsDeep Speech learns a function that is robust to background noise, reverberation and speaker variation.
Use few labelswav2vec 2.0 learns representations from speech audio before fine-tuning on transcribed speech.
3 · How it works

Follow one utterance through a Listen, Attend and Spell style recognizer.

The listener encodes filter-bank spectra; the attention-based speller emits characters.
  1. 1 · measureSupply filter-bank spectra as input to the listener.
  2. 2 · listenEncode those features with a pyramidal recurrent listener.
  3. 3 · attendUse an attention-based recurrent decoder.
  4. 4 · spellEmit the transcript as a sequence of characters.

Finding alignment is a key challenge in sequence transduction.

4 · Where it's used
WhoWhat they askWhat it works with
Captioning team“What words were spoken in this recording?”Audio and transcript output
Meeting tool builder“Where did the recognizer lose a phrase?”Acoustic frames and decoded text
Evaluation team“How many word insertions, deletions and substitutions occurred?”Reference and predicted transcripts
5 · What it solves, and what it doesn't
solves
  • An end-to-end recognizer can learn its speech-recognition components jointly.
  • Listen, Attend and Spell emits characters from filter-bank spectra.
  • Whisper was trained on multilingual and multitask supervision.
  • wav2vec 2.0 pretrains on speech audio before fine-tuning on transcripts.
doesn't solve
  • Input-output alignment remains a central sequence-transduction challenge.
  • wav2vec 2.0 still fine-tunes on transcribed speech.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperDeep Speech: Scaling up end-to-end speech recognition, Hannun and colleagues · read 28 Sept 2026
  2. paperListen, Attend and Spell, Chan and colleagues · read 28 Sept 2026
  3. paperSequence Transduction with Recurrent Neural Networks, Graves · read 28 Sept 2026
  4. paperRobust Speech Recognition via Large-Scale Weak Supervision, Radford and colleagues · read 28 Sept 2026
  5. paperwav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations, Baevski and colleagues · read 28 Sept 2026