Concepts

Vision-language-action models

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

A vision-language-action model maps visual observations and a language instruction to robot actions.

1 · What it is

A vision-language-action model, or VLA, extends multimodal understanding into control. It consumes visual observations and an instruction, then predicts actions a robot can execute. RT-2, for example, represents actions as text tokens.

RT-2 established a direct recipe: encode robot actions as text tokens and train those tokens alongside Internet-scale vision-language tasks. OpenVLA was trained on 970,000 real-world robot demonstrations.

Other designs use different action representations. Pi-zero builds a flow-matching architecture on a pretrained vision-language model. Octo can take language commands or goal images and can be fine-tuned for new sensory inputs and action spaces.

Vision-language pretraining is not enough by itself. RT-2 combines Internet-scale vision-language tasks with robot trajectory data. Octo’s results also show that new observation and action spaces can require fine-tuning.

2 · Why it exists

A robot policy must connect visual and linguistic context to actions in the physical world.

Ground languageThe instruction has to refer to objects and relations in the current scene.
Produce controlThe output must represent robot actions rather than only natural-language tokens.
Reuse knowledgeA pretrained vision-language model can provide a starting point for robot action training.
3 · How it works

Encode images, robot state and the instruction, then decode an action or action chunk.

A VLA closes the gap between multimodal understanding and robot control.
  1. 1 · observeCollect visual observations and the robot's current state.
  2. 2 · instructSupply the task as a natural-language instruction or goal image.
  3. 3 · predictDecode the model representation into robot actions.
  4. 4 · executeUse the predicted controls for end-to-end robot control.

RT-2 writes actions as text tokens. Pi-zero builds a flow-matching architecture on a pretrained vision-language model.

4 · Where it's used
WhoWhat they askWhat it works with
Manipulation robot“Put the named object in the bowl.”Camera views, instruction and robot state
Mobile manipulator“Carry laundry from the dryer to the table.”Visual observations across the task
Robot researcher“Adapt one policy to a new robot setup.”Demonstrations from the new sensors and action space
Generalist policy“Follow either a language command or a goal image.”The chosen goal representation
5 · What it solves, and what it doesn't
solves
  • RT-2 expresses robot actions as text tokens and trains them alongside vision-language tasks.
  • OpenVLA was trained on 970,000 real-world robot demonstrations.
  • Octo accepts language commands or goal images and can be fine-tuned to new sensor and action spaces.
  • pi-zero builds a flow-matching architecture on a pretrained vision-language model.
doesn't solve
  • Vision-language pretraining does not remove the need for robot trajectory data.
  • A policy trained on particular robots may need adaptation for a new sensor or action space.
  • Robot learning still faces obstacles in data, generalization and robustness.
  • Grounding language in real-world continuous sensor inputs remains a central robotics challenge.