Speech-to-text models
A speech-to-text model can transcribe a speech utterance into written characters.
Speech-to-text turns audio into a written sequence. In sequence transduction, the alignment between discrete input and output sequences can be unknown.
Listen, Attend and Spell uses two components. A pyramidal recurrent listener accepts filter-bank spectra. An attention-based recurrent speller emits characters. The complete recognizer learns its components jointly.
Other systems use different training routes. Deep Speech used end-to-end deep learning without a phoneme dictionary. Whisper training used multilingual and multitask supervision. wav2vec 2.0 masked speech input in latent space and solved a contrastive task before transcript fine-tuning.
Speech transduction can involve an unknown input-output alignment.
Follow one utterance through a Listen, Attend and Spell style recognizer.
- 1 · measureSupply filter-bank spectra as input to the listener.
- 2 · listenEncode those features with a pyramidal recurrent listener.
- 3 · attendUse an attention-based recurrent decoder.
- 4 · spellEmit the transcript as a sequence of characters.
Finding alignment is a key challenge in sequence transduction.
| Who | What they ask | What it works with |
|---|---|---|
| Captioning team | “What words were spoken in this recording?” | Audio and transcript output |
| Meeting tool builder | “Where did the recognizer lose a phrase?” | Acoustic frames and decoded text |
| Evaluation team | “How many word insertions, deletions and substitutions occurred?” | Reference and predicted transcripts |
- An end-to-end recognizer can learn its speech-recognition components jointly.
- Listen, Attend and Spell emits characters from filter-bank spectra.
- Whisper was trained on multilingual and multitask supervision.
- wav2vec 2.0 pretrains on speech audio before fine-tuning on transcripts.
- Input-output alignment remains a central sequence-transduction challenge.
- wav2vec 2.0 still fine-tunes on transcribed speech.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperDeep Speech: Scaling up end-to-end speech recognition, Hannun and colleagues · read 28 Sept 2026
- paperListen, Attend and Spell, Chan and colleagues · read 28 Sept 2026
- paperSequence Transduction with Recurrent Neural Networks, Graves · read 28 Sept 2026
- paperRobust Speech Recognition via Large-Scale Weak Supervision, Radford and colleagues · read 28 Sept 2026
- paperwav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations, Baevski and colleagues · read 28 Sept 2026