Text-to-image models
A text-to-image model can use text as a conditioning input to an image generator.
Text-to-image systems can use text as a conditioning input to an image generator. Imagen encodes that text with a pretrained text-only language model.
One route is latent diffusion. Cross-attention carries conditioning into a diffusion model in the latent space of a pretrained autoencoder. The final latent sample can be decoded to image space in one pass.
Other routes use image tokens. DALL-E modeled text and image tokens as one autoregressive stream. Parti treated text-to-image generation as sequence-to-sequence modeling with image tokens as outputs. Muse instead used masked modeling in discrete token space.
Text can act as a conditioning input to an image generator.
Follow one prompt through a latent-diffusion style path.
- 1 · encodeConvert the prompt into a text representation.
- 2 · conditionFeed that representation into cross-attention layers in the generator.
- 3 · generateRun the diffusion model in the latent space of a pretrained autoencoder.
- 4 · decodeDecode the final latent sample to image space in one decoder pass.
Muse is one alternative design: it uses masked modeling in discrete token space.
| Who | What they ask | What it works with |
|---|---|---|
| Creative tool builder | “Does this prompt produce the requested composition?” | Prompt and generated image pairs |
| Model researcher | “Should the generator work in pixels, latents or image tokens?” | The model's image representation |
| Evaluation team | “Does the output match the text?” | Image-text alignment results |
- Text conditioning can guide a diffusion generator through cross-attention.
- A latent diffusion model can run in the latent space of a pretrained autoencoder.
- An autoregressive model can treat text and image tokens as one stream.
- A masked generative model can predict masked image tokens from a text embedding.
- Pixel-space diffusion can require sequential evaluations at inference.
- An image-token route requires a tokenizer that encodes images as discrete tokens.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperZero-Shot Text-to-Image Generation, Ramesh and colleagues · read 28 Sept 2026
- paperPhotorealistic Text-to-Image Diffusion Models with Deep Language Understanding, Saharia and colleagues · read 28 Sept 2026
- paperHigh-Resolution Image Synthesis with Latent Diffusion Models, Rombach and colleagues · read 28 Sept 2026
- paperScaling Autoregressive Models for Content-Rich Text-to-Image Generation, Yu and colleagues · read 28 Sept 2026
- paperMuse: Text-To-Image Generation via Masked Generative Transformers, Chang and colleagues · read 28 Sept 2026