Concepts

Text-to-image models

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

A text-to-image model can use text as a conditioning input to an image generator.

1 · What it is

Text-to-image systems can use text as a conditioning input to an image generator. Imagen encodes that text with a pretrained text-only language model.

One route is latent diffusion. Cross-attention carries conditioning into a diffusion model in the latent space of a pretrained autoencoder. The final latent sample can be decoded to image space in one pass.

Other routes use image tokens. DALL-E modeled text and image tokens as one autoregressive stream. Parti treated text-to-image generation as sequence-to-sequence modeling with image tokens as outputs. Muse instead used masked modeling in discrete token space.

2 · Why it exists

Text can act as a conditioning input to an image generator.

Read the promptImagen uses a pretrained text-only language model to encode text for image synthesis.
Build the imageAutoregressive systems can represent an image as a sequence of discrete image tokens.
Control generationLatent diffusion models use cross-attention for conditioning inputs such as text.
3 · How it works

Follow one prompt through a latent-diffusion style path.

The final latent sample can be decoded to image space in one pass.
  1. 1 · encodeConvert the prompt into a text representation.
  2. 2 · conditionFeed that representation into cross-attention layers in the generator.
  3. 3 · generateRun the diffusion model in the latent space of a pretrained autoencoder.
  4. 4 · decodeDecode the final latent sample to image space in one decoder pass.

Muse is one alternative design: it uses masked modeling in discrete token space.

4 · Where it's used
WhoWhat they askWhat it works with
Creative tool builder“Does this prompt produce the requested composition?”Prompt and generated image pairs
Model researcher“Should the generator work in pixels, latents or image tokens?”The model's image representation
Evaluation team“Does the output match the text?”Image-text alignment results
5 · What it solves, and what it doesn't
solves
  • Text conditioning can guide a diffusion generator through cross-attention.
  • A latent diffusion model can run in the latent space of a pretrained autoencoder.
  • An autoregressive model can treat text and image tokens as one stream.
  • A masked generative model can predict masked image tokens from a text embedding.
doesn't solve
  • Pixel-space diffusion can require sequential evaluations at inference.
  • An image-token route requires a tokenizer that encodes images as discrete tokens.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperZero-Shot Text-to-Image Generation, Ramesh and colleagues · read 28 Sept 2026
  2. paperPhotorealistic Text-to-Image Diffusion Models with Deep Language Understanding, Saharia and colleagues · read 28 Sept 2026
  3. paperHigh-Resolution Image Synthesis with Latent Diffusion Models, Rombach and colleagues · read 28 Sept 2026
  4. paperScaling Autoregressive Models for Content-Rich Text-to-Image Generation, Yu and colleagues · read 28 Sept 2026
  5. paperMuse: Text-To-Image Generation via Masked Generative Transformers, Chang and colleagues · read 28 Sept 2026