Encoder-decoderConcepts
Encoder-decoder architecture
1 · In one line
An encoder-decoder model combines an encoder with an autoregressive decoder for sequence generation.
1 · What it is
EncoderDecoderModel can be initialized with an encoder and a decoder. The encoder maps the input to continuous representations. The decoder produces the next output element.
The Transformer uses stacked self-attention in both halves. Its decoder also attends to the encoder stack. Cross-attention lets each decoder position read the encoder output.
BART uses a Transformer-based sequence-to-sequence architecture. T5 casts language problems into a text-to-text format.
Many tasks receive one sequence and must produce another sequence.
Different lengthsA translation can contain a different number of tokens from its source sentence.
Source accessAttention lets the decoder search source positions that matter for its next output.
Output orderAn autoregressive decoder generates the output one element at a time.
Follow one sentence from source tokens to translated tokens.
- 1 · encodeThe encoder maps the input sequence to a sequence of continuous representations.
- 2 · startThe decoder receives the output tokens already available to it.
- 3 · attendEncoder-decoder attention reads keys and values from the encoder output.
- 4 · predictThe decoder produces the next output element.
The decoder uses cross-attention over the encoder output.
| Who | What they ask | What it works with |
|---|---|---|
| Translation system | “How should this source sentence be written in French?” | Encoded source positions and earlier translated tokens |
| Summarization system | “What short sequence represents this document?” | Encoded document positions and the summary prefix |
| Speech recognizer | “Which text sequence corresponds to this audio?” | Encoded input features and earlier output tokens |
solves
- The encoder and decoder can process different sequence lengths.
- Cross-attention lets each decoder position read the encoder output.
- Text-to-text models can use one format across multiple language tasks.
doesn't solve
- A fixed-length encoder summary can become an information bottleneck.
- Encoder-decoder models require downstream fine-tuning for a specific task.
6 · Go deeper
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperAttention Is All You Need, Vaswani et al. · read 28 Sept 2026
- paperNeural Machine Translation by Jointly Learning to Align and Translate, Bahdanau, Cho and Bengio · read 28 Sept 2026
- paperExploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, Raffel et al. · read 28 Sept 2026
- paperBART Denoising Sequence-to-Sequence Pre-training for Natural Language Generation Translation and Comprehension, Lewis et al. · read 28 Sept 2026
- docsEncoder Decoder Models, Hugging Face · read 28 Sept 2026