World models
A world model predicts how an environment may change after an action, so an agent can evaluate possible futures before acting.
A world model is a learned simulator. It receives a representation of the current environment and a candidate action, then predicts information about the state that may follow. A planner or policy can compare those imagined futures before committing an action to the real environment.
The prediction does not have to be a perfect video. The original World Models work learned a compressed spatial and temporal representation. PlaNet planned in a learned latent space. MuZero went further by predicting only the information its planner required: reward, policy and value.
World models also support learning from imagined experience. Dreamer propagates gradients through latent trajectories generated by its learned model. The original World Models agent could train a policy inside generated experience and transfer that policy back to the actual environment.
Generative world models widen the setting. Genie produces action-controllable virtual worlds from text, synthetic images, photographs and sketches. Accurate learned dynamics remain a core requirement: PlaNet’s authors describe models accurate enough for planning as a long-standing challenge.
Acting only in the real environment can make learning slow, expensive or unsafe.
Encode the present, apply a candidate action in latent space, then score the predicted future.
- 1 · encodeCompress the current observation into a latent state.
- 2 · imagineApply candidate actions to the learned dynamics model.
- 3 · scorePredict rewards, values or observations along the imagined trajectories.
- 4 · actChoose an action and execute it in the real environment.
A useful world model need not reproduce every pixel. MuZero learns the parts of the environment needed for planning: reward, policy and value.
| Who | What they ask | What it works with |
|---|---|---|
| Robot learner | “What may happen if the gripper moves left?” | Predicted future states under candidate controls |
| Game agent | “Which move has the best simulated continuation?” | Imagined rewards and values |
| Control system | “Which action sequence reaches the target state?” | Rollouts in a learned latent dynamics model |
| Environment researcher | “Generate new interactive worlds for an agent.” | Action-conditioned predicted frames |
- The original World Models system learned a compressed spatial and temporal representation of its environment.
- Dreamer learns behavior by backpropagating through trajectories imagined in a compact latent state space.
- MuZero plans with a learned model that predicts reward, policy and value.
- Genie can generate action-controllable virtual worlds from text, images, photographs or sketches.
- Learning dynamics that are accurate enough for planning remains a long-standing challenge.
- Multi-step planning requires the dynamics model to predict rewards accurately several steps ahead.
- A world model does not remove the planner or policy that chooses actions.
- A policy trained in generated experience still has to transfer back to the actual environment.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperWorld Models, Ha and Schmidhuber · read 28 Sept 2026
- paperDream to Control: Learning Behaviors by Latent Imagination, Hafner et al. · read 28 Sept 2026
- paperMastering Atari, Go, Chess and Shogi by Planning with a Learned Model, Schrittwieser et al. · read 28 Sept 2026
- paperGenie: Generative Interactive Environments, Bruce et al. · read 28 Sept 2026
- paperLearning Latent Dynamics for Planning from Pixels, Hafner et al. · read 28 Sept 2026