Concepts

World models

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

A world model predicts how an environment may change after an action, so an agent can evaluate possible futures before acting.

1 · What it is

A world model is a learned simulator. It receives a representation of the current environment and a candidate action, then predicts information about the state that may follow. A planner or policy can compare those imagined futures before committing an action to the real environment.

The prediction does not have to be a perfect video. The original World Models work learned a compressed spatial and temporal representation. PlaNet planned in a learned latent space. MuZero went further by predicting only the information its planner required: reward, policy and value.

World models also support learning from imagined experience. Dreamer propagates gradients through latent trajectories generated by its learned model. The original World Models agent could train a policy inside generated experience and transfer that policy back to the actual environment.

Generative world models widen the setting. Genie produces action-controllable virtual worlds from text, synthetic images, photographs and sketches. Accurate learned dynamics remain a core requirement: PlaNet’s authors describe models accurate enough for planning as a long-standing challenge.

2 · Why it exists

Acting only in the real environment can make learning slow, expensive or unsafe.

Predict changeA learned dynamics model can estimate a future latent state from a current state and action.
Imagine trajectoriesAn agent can train or plan over sequences generated inside the model.
Compress experienceThe model can represent high-dimensional observations in a smaller latent state.
3 · How it works

Encode the present, apply a candidate action in latent space, then score the predicted future.

The world model supplies predicted consequences; a planner or policy still chooses the action.
  1. 1 · encodeCompress the current observation into a latent state.
  2. 2 · imagineApply candidate actions to the learned dynamics model.
  3. 3 · scorePredict rewards, values or observations along the imagined trajectories.
  4. 4 · actChoose an action and execute it in the real environment.

A useful world model need not reproduce every pixel. MuZero learns the parts of the environment needed for planning: reward, policy and value.

4 · Where it's used
WhoWhat they askWhat it works with
Robot learner“What may happen if the gripper moves left?”Predicted future states under candidate controls
Game agent“Which move has the best simulated continuation?”Imagined rewards and values
Control system“Which action sequence reaches the target state?”Rollouts in a learned latent dynamics model
Environment researcher“Generate new interactive worlds for an agent.”Action-conditioned predicted frames
5 · What it solves, and what it doesn't
solves
  • The original World Models system learned a compressed spatial and temporal representation of its environment.
  • Dreamer learns behavior by backpropagating through trajectories imagined in a compact latent state space.
  • MuZero plans with a learned model that predicts reward, policy and value.
  • Genie can generate action-controllable virtual worlds from text, images, photographs or sketches.
doesn't solve
  • Learning dynamics that are accurate enough for planning remains a long-standing challenge.
  • Multi-step planning requires the dynamics model to predict rewards accurately several steps ahead.
  • A world model does not remove the planner or policy that chooses actions.
  • A policy trained in generated experience still has to transfer back to the actual environment.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperWorld Models, Ha and Schmidhuber · read 28 Sept 2026
  2. paperDream to Control: Learning Behaviors by Latent Imagination, Hafner et al. · read 28 Sept 2026
  3. paperMastering Atari, Go, Chess and Shogi by Planning with a Learned Model, Schrittwieser et al. · read 28 Sept 2026
  4. paperGenie: Generative Interactive Environments, Bruce et al. · read 28 Sept 2026
  5. paperLearning Latent Dynamics for Planning from Pixels, Hafner et al. · read 28 Sept 2026