Concepts

Text-to-video models

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

Given a text prompt, a text-to-video model generates a video.

1 · What it is

Text-to-video models generate video from text conditioning. Make-A-Video learns appearance from paired text-image data and motion from unlabelled video footage.

Imagen Video uses a base video generation model followed by interleaved spatial and temporal video super-resolution models. Lumiere instead generates the full temporal duration in one model pass.

Video Diffusion Models extended a standard image diffusion architecture to video generation. VideoPoet used a decoder-only Transformer over multimodal inputs. Make-A-Video separated learning appearance from paired text-image data and learning motion from unlabelled video.

2 · Why it exists

Make-A-Video separates learning appearance from learning motion.

Visual contentMake-A-Video learns appearance and descriptions from paired text-image data.
MotionMake-A-Video learns how the world moves from unlabelled video footage.
Temporal coherenceVideo diffusion research identifies temporally coherent high-fidelity video as a central goal.
3 · How it works

Follow one prompt through an Imagen Video style cascade.

Imagen Video first generates a base video, then applies interleaved spatial and temporal video super-resolution models.
  1. 1 · promptBegin with the text prompt that conditions video generation.
  2. 2 · generateUse a base video generation model.
  3. 3 · refineSpatial and temporal super-resolution models refine the video in sequence.
  4. 4 · returnThe cascade produces a high-definition video.

Lumiere generates the full temporal duration in one model pass.

4 · Where it's used
WhoWhat they askWhat it works with
Video tool builder“Does the clip follow the requested action?”Prompt and generated clip pairs
Model researcher“Is time generated jointly or refined in later stages?”The temporal architecture
Evaluation team“Does motion remain coherent across the clip?”Temporal consistency results
5 · What it solves, and what it doesn't
solves
  • Text prompts can condition a video diffusion system.
  • A cascade can separate base generation from spatial and temporal super-resolution.
  • Lumiere generates the full temporal duration in one model pass.
doesn't solve
  • Training large-scale text-to-video foundation models remains an open challenge because motion adds complexity.
  • The temporal data dimension creates memory, compute and training-data challenges.
  • Existing models remain restricted in video duration, visual quality and realistic motion.
  • A high-resolution super-resolution network cannot process Lumiere's entire video duration within its memory requirements.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperMake-A-Video: Text-to-Video Generation without Text-Video Data, Singer and colleagues · read 28 Sept 2026
  2. paperImagen Video: High Definition Video Generation with Diffusion Models, Ho and colleagues · read 28 Sept 2026
  3. paperVideo Diffusion Models, Ho and colleagues · read 28 Sept 2026
  4. paperLumiere: A Space-Time Diffusion Model for Video Generation, Bar-Tal and colleagues · read 28 Sept 2026
  5. paperVideoPoet: A Large Language Model for Zero-Shot Video Generation, Kondratyuk and colleagues · read 28 Sept 2026