Encoder-decoderConcepts

Encoder-decoder architecture

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

An encoder-decoder model combines an encoder with an autoregressive decoder for sequence generation.

1 · What it is

EncoderDecoderModel can be initialized with an encoder and a decoder. The encoder maps the input to continuous representations. The decoder produces the next output element.

The Transformer uses stacked self-attention in both halves. Its decoder also attends to the encoder stack. Cross-attention lets each decoder position read the encoder output.

BART uses a Transformer-based sequence-to-sequence architecture. T5 casts language problems into a text-to-text format.

2 · Why it exists

Many tasks receive one sequence and must produce another sequence.

Different lengthsA translation can contain a different number of tokens from its source sentence.
Source accessAttention lets the decoder search source positions that matter for its next output.
Output orderAn autoregressive decoder generates the output one element at a time.
3 · How it works

Follow one sentence from source tokens to translated tokens.

The encoder maps the source to continuous representations. The decoder attends to the encoder output while generating the target sequence.
  1. 1 · encodeThe encoder maps the input sequence to a sequence of continuous representations.
  2. 2 · startThe decoder receives the output tokens already available to it.
  3. 3 · attendEncoder-decoder attention reads keys and values from the encoder output.
  4. 4 · predictThe decoder produces the next output element.

The decoder uses cross-attention over the encoder output.

4 · Where it's used
WhoWhat they askWhat it works with
Translation system“How should this source sentence be written in French?”Encoded source positions and earlier translated tokens
Summarization system“What short sequence represents this document?”Encoded document positions and the summary prefix
Speech recognizer“Which text sequence corresponds to this audio?”Encoded input features and earlier output tokens
5 · What it solves, and what it doesn't
solves
  • The encoder and decoder can process different sequence lengths.
  • Cross-attention lets each decoder position read the encoder output.
  • Text-to-text models can use one format across multiple language tasks.
doesn't solve
  • A fixed-length encoder summary can become an information bottleneck.
  • Encoder-decoder models require downstream fine-tuning for a specific task.