Text-to-video models
Given a text prompt, a text-to-video model generates a video.
Text-to-video models generate video from text conditioning. Make-A-Video learns appearance from paired text-image data and motion from unlabelled video footage.
Imagen Video uses a base video generation model followed by interleaved spatial and temporal video super-resolution models. Lumiere instead generates the full temporal duration in one model pass.
Video Diffusion Models extended a standard image diffusion architecture to video generation. VideoPoet used a decoder-only Transformer over multimodal inputs. Make-A-Video separated learning appearance from paired text-image data and learning motion from unlabelled video.
Make-A-Video separates learning appearance from learning motion.
Follow one prompt through an Imagen Video style cascade.
- 1 · promptBegin with the text prompt that conditions video generation.
- 2 · generateUse a base video generation model.
- 3 · refineSpatial and temporal super-resolution models refine the video in sequence.
- 4 · returnThe cascade produces a high-definition video.
Lumiere generates the full temporal duration in one model pass.
| Who | What they ask | What it works with |
|---|---|---|
| Video tool builder | “Does the clip follow the requested action?” | Prompt and generated clip pairs |
| Model researcher | “Is time generated jointly or refined in later stages?” | The temporal architecture |
| Evaluation team | “Does motion remain coherent across the clip?” | Temporal consistency results |
- Text prompts can condition a video diffusion system.
- A cascade can separate base generation from spatial and temporal super-resolution.
- Lumiere generates the full temporal duration in one model pass.
- Training large-scale text-to-video foundation models remains an open challenge because motion adds complexity.
- The temporal data dimension creates memory, compute and training-data challenges.
- Existing models remain restricted in video duration, visual quality and realistic motion.
- A high-resolution super-resolution network cannot process Lumiere's entire video duration within its memory requirements.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperMake-A-Video: Text-to-Video Generation without Text-Video Data, Singer and colleagues · read 28 Sept 2026
- paperImagen Video: High Definition Video Generation with Diffusion Models, Ho and colleagues · read 28 Sept 2026
- paperVideo Diffusion Models, Ho and colleagues · read 28 Sept 2026
- paperLumiere: A Space-Time Diffusion Model for Video Generation, Bar-Tal and colleagues · read 28 Sept 2026
- paperVideoPoet: A Large Language Model for Zero-Shot Video Generation, Kondratyuk and colleagues · read 28 Sept 2026